ModelRefs / Evaluation Pipeline — Canonical Workflow
Evaluation Pipeline — Canonical Workflow
Evaluation Pipeline: provisional AI workflow implementation reference with candidate models, providers, tools, and architecture.
Overview
An evaluation pipeline is a continuous offline/online harness for testing model, prompt, or system changes against a baseline — regression detection, dataset versioning, and dashboards so quality, cost, or latency regressions surface before reaching production.
Use this page to check what an evaluation harness needs to cover for your workload — baseline comparison, dataset versioning, and regression thresholds — then review the related architecture, tool, and benchmark references before implementing one.
This workflow's representative-workload evaluation protocol has been drafted, but ModelRefs does not yet have recorded run evidence against it. Treat readiness or completion claims as provisional and evaluate the harness on your own representative workload before relying on it for production gating.
Implementation profile
| Category | llms |
|---|---|
| Implementation maturity | enterprise |
| Evidence status | partial |
| Primary use cases | reasoning |
| Deployment options | managed-api, hybrid |
| Architectures | serverless-api, managed-container, self-hosted-cluster |
Candidate models with published references
- BGE-M3
- GPT-5
- GPT-5 Mini
- Claude Opus 4
- Llama 4 Scout
- DeepSeek R1
- Mistral Large 2
- Command R+
- o3
- o4 Mini
- Text Embedding 3 Large
- Claude Sonnet 4
Coverage means the model is a candidate worth evaluating for this workflow, not a ranking or a recommendation. Models whose reference pages are still in review are omitted.
Benchmarks relevant to this workflow
miracl, mkqa, mldr, swe-bench, aider-polyglot, gpqa, aime-2025, tau-bench, browsecomp-long-context, longfact-concepts, terminal-bench, mmmu, mmlu-pro, livecodebench.
Relevance is a coverage signal from the canonical registry. Each benchmark only describes its own protocol and date, so confirm the harness matches your workload before treating a score as evidence.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Evaluation Pipeline — Canonical Workflow.