Аудит целостности ML-экспериментов
★ 7.1 · testing
experiment-audit is a Claude Code skill that verifies experiment integrity before results are claimed, using cross-model review through an external reviewer backend. The skill enforces a strict separation of roles: the Claude executor collects only file paths — evaluation scripts, result files, ground truth references, experiment tracker, paper drafts, and configs — while the independent reviewer (Codex MCP or Manual Review MCP) reads the code directly and delivers the verdict. Four fraud patterns are checked: fake ground truth constructed from model outputs, score normalization against the model's own maximum, phantom results citing files never created or functions never called, and insufficient evaluation scope. Designed for ML researchers and engineers who need to confirm evaluation honesty before writing claims or submitting papers.
- #experiment-integrity
- #result-validation
- #cross-model-review
- #fraud-detection
- #quality-assurance