NeMo Evaluator — оценка LLM на 100+ бенчмарках
★ 7.7 · testing
nemo-evaluator is a Claude Code skill that runs LLM benchmarking across 100+ tasks from 18+ harnesses, covering MMLU, HumanEval, GSM8K, GPQA Diamond, VLM, and safety evaluations. Built on NVIDIA's NeMo Evaluator SDK, it supports three execution backends — local Docker, Slurm HPC, and Lepton cloud — with a container-first design for reproducible results. Workflows are configured via Hydra YAML files and controlled through `nemo-evaluator-launcher` CLI commands: `run`, `status`, `ls`, `info`, `kill`, and `export` to MLflow. Ideal for ML engineers who need to compare multiple models on identical benchmarks or scale evaluation to HPC clusters with configurable tensor and data parallelism.
- #nemo
- #nvidia
- #benchmarking
- #slurm
- #docker
- #distributed-evaluation