l2-bench-pipeline / README.md
jimmyedgell's picture
Update README.md
0e1984a verified
|
Raw History Blame Contribute Delete
7.85 kB
metadata
title: L2-Bench Pipeline
emoji: πŸ“Š
colorFrom: blue
colorTo: indigo
license: mit
sdk: static
app_file: index.html
pinned: false
short_description: Reference evaluation implementation for L2-Bench
datasets:
  - OUP/l2-bench

L2-Bench Pipeline is the reference evaluation implementation for L2-Bench, including response generation and LLM-as-a-Judge scoring pipeline with inspect-ai.

Note that we use a slightly modified internal evaluation implementation to this when we report our L2-Bench results and experiments that is specific to our internal platforms, but we release this official external implementation to lower the barrier for researchers and practitioners to use L2-Bench.

Resources

Quick Start

Prerequisites

  1. Python 3.13+ with uv.
  2. Model credentials in a repo-root .env file.
  3. Dataset access. The L2-Bench dataset is gated with automatic approval. Once, in a browser, visit OUP/l2-bench while signed in and accept the terms β€” approval is instant. This cannot be done from the CLI.
  4. A token on this machine: hf auth login, or set HF_TOKEN in .env.
git clone https://huggingface.co/spaces/OUP/l2-bench-pipeline
cd l2-bench-pipeline

uv venv
source .venv/bin/activate
uv pip install -e "."

# Run the benchmark (solver + scorer)
uv run python -m l2_bench_eval.eval \
    --model bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0 \
    --log-dir logs/my-eval-run

# Smoke test with 2 samples
uv run python -m l2_bench_eval.eval \
    --model bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0 \
    --log-dir logs/smoke-test \
    --sample-limit 2

# View results
inspect view --log-dir logs/

Methodology

This pipeline implements the measurement procedure described in the paper. It is built using inspect-ai: a model under test (the solver) generates a response to each task, then an LLM-as-judge (the scorer) returns an independent binary Pass/Fail verdict for every criterion, which are combined by weighted aggregation.

The production judge is Claude Sonnet 4.6 via AWS Bedrock, judge prompt v1, a 1,024-token reasoning budget, and no verdict retries β€” the configuration selected in the paper (Appendix H). These are the pipeline defaults, declared in src/l2_bench_eval/config.py.

Note that judge temperature is left unset, since extended thinking and an explicit temperature are mutually exclusive for Claude on Bedrock, so the provider default applies whenever a reasoning budget is in use. Solver temperature is 0.0 by default.

N.B. Versions v2–v4 are retained in this repo for reproducibility (they are the 2Γ—2 prompt ablation reported in the paper Appendix H).

Tasks and resources are fetched from the dataset repo and cached locally, so a clean checkout runs without any local data files. The dataset is gated, so this fetch requires that you have accepted its terms and have a token available (prerequisites 3 and 4 above).

However, passing both --csv-path and --resources-dir bypasses the Hugging Face fetch and uses local files instead.

Key parameters

Parameter Default Description
--model (required) Solver model id (e.g. bedrock/..., openai/...)
--log-dir (required) Directory for output .eval logs
--epochs 1 Number of evaluation epochs
--sample-limit 0 (all) Limit to the first N tasks (0 = all 1,000)
--task-ids all Evaluate only these task_key values
--solver-temperature 0.0 Solver sampling temperature
--solver-max-tokens 4096 Solver max output tokens
--scorer-model bedrock/us.anthropic.claude-sonnet-4-6 Judge model
--scorer-reasoning-tokens 1024 Judge reasoning budget
--scorer-temperature unset Judge sampling temperature (see note above)
--prompt-version v1 Judge prompt version
--judge-verdict-retries 0 Re-prompts when the judge returns an unparseable verdict
--continue-on-fail True Keep going after transient errors

Every --solver-* and --scorer-* flag accepted by inspect-ai's GenerateConfig is exposed; run with --help for the full list. Note that --scorer-max-retries is the judge's API retry count, distinct from --judge-verdict-retries.

Task data

Passing both --csv-path and --resources-dir bypasses the Hugging Face fetch and uses local files instead. Otherwise these settings apply, all overridable by environment variable (e.g. in .env):

Variable Default Description
TASKS_REPO OUP/l2-bench HF dataset repo id for tasks + resources
TASKS_CSV_FILE l2-bench_tasks.csv Task CSV filename within that repo
TASKS_RESOURCES_DIR resources_for_tasks Resources directory within that repo
DEFAULT_JUDGE_MODEL bedrock/us.anthropic.claude-sonnet-4-6 Production judge model
DEFAULT_JUDGE_PROMPT_VERSION v1 Production judge prompt
DEFAULT_JUDGE_REASONING_TOKENS 1024 Production judge reasoning budget
DEFAULT_JUDGE_MAX_RETRIES 0 Production judge verdict retries

Source Files

File Description
index.html This Space's static landing page only (not needed for pipeline)
src/l2_bench_eval/eval.py CLI entry point (solver + scorer)
src/l2_bench_eval/config.py Dataset location and production judge defaults
src/l2_bench_eval/task.py Builds the inspect-ai Task; fetches data from the Hub
src/l2_bench_eval/dataset.py Parses l2-bench_tasks.csv + resources into a MemoryDataset
src/l2_bench_eval/criteria.py Rubric parsing and per-task criterion lookup
src/l2_bench_eval/score.py LLM-as-judge scorer and weighted aggregation
src/l2_bench_eval/prompts/ Judge prompt templates (v1–v4) and retry suffix
src/l2_bench_eval/bedrock_patch.py Raises the Bedrock client read timeout for long responses

Licence

This reference implementation pipeline is released under the MIT Licence. Β© 2026 Oxford University Press.

The L2-Bench dataset is licensed separately under CC-BY-SA-4.0; see the dataset card.

Citation

If you use L2-Bench in your research, please cite our work. The full citation is available below:

@misc{edgell2026l2benchevaluationbenchmarkmeasuring,
      title={L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education}, 
      author={James Edgell and Wm. Matthew Kennedy and Ben Knight and Danielle Carvalho and Martin Ku and Isaac Pattis},
      year={2026},
      eprint={2607.08842},
      archivePrefix={arXiv},
      primaryClass={cs.CY},
      url={https://arxiv.org/abs/2607.08842}, 
}

For any inquiries or feedback, including submitting corrections, please use the "Register Interest" webform on our L2-Bench site and we will respond as soon as we can.

Version History

Version Date Author Changes
1.0.0 2026-07-30 M. Ku Initial open-source release