ForgeRL / Agent research
The evidence behind the router.
Recorded coding-agent experiments. Compare policies, inspect the patch, follow every decision.
Last benchmark run: —
- Catalog
- —
- Recorded episodes
- —
- Models evaluated
- —
- Benchmark version
- —
Cost / quality
Success versus accounted cost
Loading recorded measurements…
Policy comparison
Five policies. One protocol.
Missing measurements remain blank.
| Policy | Episodes | Success | Hidden tests | Cost / task | Tokens / task | Latency | Calls | Attempts | Regressions | After repair | Escalations |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Loading recorded measurements… | |||||||||||
Mean metrics are per recorded episode. Hidden-test pass rate and task success are different measures. “After repair” is conditional on repair attempts.
Reproducible tasks
Explore the benchmark
Task brief
Select an authored task
Source files and the success criterion will appear here.
Success criterion
—
No source loaded.
Source is supplied as context to the model; this is not a claim of autonomous file discovery.
Visible checks —
- No checks loaded.
Stored execution
Inspect a recorded trajectory
Select a task to see its available experiment runs.
- Hidden tests
- —
- Accounted cost
- —
- Tokens
- —
- Latency
- —
Observable prompts, supplied files, orchestration, patches and outcomes. Private model reasoning is not recorded.
- No recorded run selected.
No recorded patch selected.
No test evaluation selected.
No recorded candidate selected.
Observed failure signals
No recorded run selected.
Labels are observable proxies, not causal explanations of a model’s reasoning. Unmeasured localization, planning or context loss remain unclassified.
Method & boundaries
A routing experiment, kept inspectable.
What the controller chooses
The policy can retry, repair, escalate, roll back or stop. The language models remain unchanged. A learned controller uses training transitions; unseen states use a declared fallback.
What the tasks establish
Authored miniature Python repositories test specific coding requirements. Multi-requirement “long-horizon” tasks are a stress category, not evidence of autonomous work on large production repositories.
What remains hidden
Only visible checks guide a repair. Hidden checks grade the final candidate after routing ends. Their inputs and reference solutions are excluded from this explorer.