ForgeRL / Agent research

The evidence behind the router.

Recorded coding-agent experiments. Compare policies, inspect the patch, follow every decision.

Loading evidence

Last benchmark run: —

Catalog
Recorded episodes
Models evaluated
Benchmark version

Cost / quality

Success versus accounted cost

Higher & left is better

Loading recorded measurements…

Points appear only for policies with measured success and accounted cost. The comparison table provides the same values.

    Policy comparison

    Five policies. One protocol.

    Missing measurements remain blank.

    Measured policy outcomes. A dash means no measurement; cost includes reported charges or conservative reservations.
    Policy Episodes Success Hidden tests Cost / task Tokens / task Latency Calls Attempts Regressions After repair Escalations
    Loading recorded measurements…

    Mean metrics are per recorded episode. Hidden-test pass rate and task success are different measures. “After repair” is conditional on repair attempts.

    Reproducible tasks

    Explore the benchmark

    Loading authored repositories

    Task brief

    Select an authored task

    Source files and the success criterion will appear here.

    Success criterion

    No source loaded.

    Source is supplied as context to the model; this is not a claim of autonomous file discovery.

    Visible checks
    • No checks loaded.

    Stored execution

    Inspect a recorded trajectory

    Select a task to see its available experiment runs.

    Not selected
    Hidden tests
    Accounted cost
    Tokens
    Latency

    Observable prompts, supplied files, orchestration, patches and outcomes. Private model reasoning is not recorded.

    1. No recorded run selected.

    Method & boundaries

    A routing experiment, kept inspectable.

    Protocol & source

    What the controller chooses

    The policy can retry, repair, escalate, roll back or stop. The language models remain unchanged. A learned controller uses training transitions; unseen states use a declared fallback.

    What the tasks establish

    Authored miniature Python repositories test specific coding requirements. Multi-requirement “long-horizon” tasks are a stress category, not evidence of autonomous work on large production repositories.

    What remains hidden

    Only visible checks guide a repair. Hidden checks grade the final candidate after routing ends. Their inputs and reference solutions are excluded from this explorer.