Prompt-engineering tournaments for improving fixed target models
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-07-10 16:30:05 +02:00
core-bench.example.yaml Add Core Bench tournament specification 2026-07-10 16:30:05 +02:00
README.md Add Core Bench tournament specification 2026-07-10 16:30:05 +02:00
SPEC.md Add Core Bench tournament specification 2026-07-10 16:30:05 +02:00

Core Bench

Core Bench measures whether one model can improve another model's performance through prompt engineering alone.

The benchmark is a tournament, not a single prompt contest. Optimizer models propose and refine system prompts over multiple rounds. A fixed target model is repeatedly evaluated on hidden task rows, and prompt lineages compete on both absolute task performance and performance gained per unit of optimizer cost.

Core Bench is intended for questions such as:

  • Can prompting make a small model reliably perform a task it is initially bad at?
  • Which optimizer model is best for the first prompt rewrite?
  • Do cheaper models become equally useful once only small refinements remain?
  • What is the highest achievable score without changing target-model weights?
  • What is the cheapest cascade that performs close to the best cascade?

Core Bench does not prescribe a task, private prompt, or dataset. Benchmark authors provide their own target model, task rows, scorer, seed prompt, and optimizer candidates.

Documents

Minimal Flow

  1. Freeze a target model and inference configuration.
  2. Split task rows into feedback, promotion, final-test, and regression sets.
  3. Score the seed prompt.
  4. Fan out zero-shot prompt proposals from multiple optimizer configurations.
  5. Evaluate every prompt by running the fixed target model locally or through a fixed endpoint.
  6. Promote an absolute-quality champion and a cost-efficient champion.
  7. Refine those lineages using scores and failures from feedback rows only.
  8. Stop when improvements flatten or the budget is exhausted.
  9. Evaluate final champions once on the sealed final-test set.
  10. Publish score-versus-cost curves, lineage diagrams, and reproducible artifacts.

Status

This repository currently contains the benchmark design specification. It does not yet contain a reference runner or task dataset.