| Filename | Latest commit message | Latest commit date |
|---|---|---|
| core-bench.example.yaml | ||
| README.md | ||
| SPEC.md | ||
Core Bench
Core Bench measures whether one model can improve another model's performance through prompt engineering alone.
The benchmark is a tournament, not a single prompt contest. Optimizer models propose and refine system prompts over multiple rounds. A fixed target model is repeatedly evaluated on hidden task rows, and prompt lineages compete on both absolute task performance and performance gained per unit of optimizer cost.
Core Bench is intended for questions such as:
- Can prompting make a small model reliably perform a task it is initially bad at?
- Which optimizer model is best for the first prompt rewrite?
- Do cheaper models become equally useful once only small refinements remain?
- What is the highest achievable score without changing target-model weights?
- What is the cheapest cascade that performs close to the best cascade?
Core Bench does not prescribe a task, private prompt, or dataset. Benchmark authors provide their own target model, task rows, scorer, seed prompt, and optimizer candidates.
Documents
- SPEC.md defines the tournament and reporting contract.
- core-bench.example.yaml shows a prompt-free setup.
Minimal Flow
- Freeze a target model and inference configuration.
- Split task rows into feedback, promotion, final-test, and regression sets.
- Score the seed prompt.
- Fan out zero-shot prompt proposals from multiple optimizer configurations.
- Evaluate every prompt by running the fixed target model locally or through a fixed endpoint.
- Promote an absolute-quality champion and a cost-efficient champion.
- Refine those lineages using scores and failures from feedback rows only.
- Stop when improvements flatten or the budget is exhausted.
- Evaluate final champions once on the sealed final-test set.
- Publish score-versus-cost curves, lineage diagrams, and reproducible artifacts.
Status
This repository currently contains the benchmark design specification. It does not yet contain a reference runner or task dataset.