Spaces:
Running
Download README.md from OUP/l2-bench-pipeline: direct link, hf CLI and curl.
- Browser
- Download file 7.85 kB
-
https://huggingface.co/spaces/OUP/l2-bench-pipeline/resolve/main/README.md
- Command line
-
hf download hf://spaces/OUP/l2-bench-pipeline/README.md
-
curl -L -o README.md https://huggingface.co/spaces/OUP/l2-bench-pipeline/resolve/main/README.md
title: L2-Bench Pipeline
emoji: π
colorFrom: blue
colorTo: indigo
license: mit
sdk: static
app_file: index.html
pinned: false
short_description: Reference evaluation implementation for L2-Bench
datasets:
- OUP/l2-bench
L2-Bench Pipeline is the reference evaluation implementation for L2-Bench, including response generation and LLM-as-a-Judge scoring pipeline with inspect-ai.
Note that we use a slightly modified internal evaluation implementation to this when we report our L2-Bench results and experiments that is specific to our internal platforms, but we release this official external implementation to lower the barrier for researchers and practitioners to use L2-Bench.
Resources
- Paper: L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
- Dataset: OUP/l2-bench
- L2-Bench site: L2-Bench release blog
Quick Start
Prerequisites
- Python 3.13+ with uv.
- Model credentials in a repo-root
.envfile. - Dataset access. The L2-Bench dataset is gated with automatic approval. Once, in a browser, visit OUP/l2-bench while signed in and accept the terms β approval is instant. This cannot be done from the CLI.
- A token on this machine:
hf auth login, or setHF_TOKENin.env.
git clone https://huggingface.co/spaces/OUP/l2-bench-pipeline
cd l2-bench-pipeline
uv venv
source .venv/bin/activate
uv pip install -e "."
# Run the benchmark (solver + scorer)
uv run python -m l2_bench_eval.eval \
--model bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0 \
--log-dir logs/my-eval-run
# Smoke test with 2 samples
uv run python -m l2_bench_eval.eval \
--model bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0 \
--log-dir logs/smoke-test \
--sample-limit 2
# View results
inspect view --log-dir logs/
Methodology
This pipeline implements the measurement procedure described in the paper. It is built using inspect-ai: a model under test (the solver) generates a response to each task, then an LLM-as-judge (the scorer) returns an independent binary Pass/Fail verdict for every criterion, which are combined by weighted aggregation.
The production judge is Claude Sonnet 4.6 via AWS Bedrock, judge prompt v1, a 1,024-token reasoning budget, and no verdict retries β the configuration selected in the paper (Appendix H). These are the pipeline defaults, declared in src/l2_bench_eval/config.py.
Note that judge temperature is left unset, since extended thinking and an explicit temperature are mutually exclusive for Claude on Bedrock, so the provider default applies whenever a reasoning budget is in use. Solver temperature is 0.0 by default.
N.B. Versions v2βv4 are retained in this repo for reproducibility (they are the 2Γ2 prompt ablation reported in the paper Appendix H).
Tasks and resources are fetched from the dataset repo and cached locally, so a clean checkout runs without any local data files. The dataset is gated, so this fetch requires that you have accepted its terms and have a token available (prerequisites 3 and 4 above).
However, passing both --csv-path and --resources-dir bypasses the Hugging Face fetch and uses local files instead.
Key parameters
| Parameter | Default | Description |
|---|---|---|
--model |
(required) | Solver model id (e.g. bedrock/..., openai/...) |
--log-dir |
(required) | Directory for output .eval logs |
--epochs |
1 |
Number of evaluation epochs |
--sample-limit |
0 (all) |
Limit to the first N tasks (0 = all 1,000) |
--task-ids |
all | Evaluate only these task_key values |
--solver-temperature |
0.0 |
Solver sampling temperature |
--solver-max-tokens |
4096 |
Solver max output tokens |
--scorer-model |
bedrock/us.anthropic.claude-sonnet-4-6 |
Judge model |
--scorer-reasoning-tokens |
1024 |
Judge reasoning budget |
--scorer-temperature |
unset | Judge sampling temperature (see note above) |
--prompt-version |
v1 |
Judge prompt version |
--judge-verdict-retries |
0 |
Re-prompts when the judge returns an unparseable verdict |
--continue-on-fail |
True |
Keep going after transient errors |
Every --solver-* and --scorer-* flag accepted by inspect-ai's
GenerateConfig is exposed; run with --help for the full list. Note that
--scorer-max-retries is the judge's API retry count, distinct from
--judge-verdict-retries.
Task data
Passing both --csv-path and --resources-dir bypasses the Hugging Face
fetch and uses local files instead. Otherwise these settings apply, all
overridable by environment variable (e.g. in .env):
| Variable | Default | Description |
|---|---|---|
TASKS_REPO |
OUP/l2-bench |
HF dataset repo id for tasks + resources |
TASKS_CSV_FILE |
l2-bench_tasks.csv |
Task CSV filename within that repo |
TASKS_RESOURCES_DIR |
resources_for_tasks |
Resources directory within that repo |
DEFAULT_JUDGE_MODEL |
bedrock/us.anthropic.claude-sonnet-4-6 |
Production judge model |
DEFAULT_JUDGE_PROMPT_VERSION |
v1 |
Production judge prompt |
DEFAULT_JUDGE_REASONING_TOKENS |
1024 |
Production judge reasoning budget |
DEFAULT_JUDGE_MAX_RETRIES |
0 |
Production judge verdict retries |
Source Files
| File | Description |
|---|---|
index.html |
This Space's static landing page only (not needed for pipeline) |
src/l2_bench_eval/eval.py |
CLI entry point (solver + scorer) |
src/l2_bench_eval/config.py |
Dataset location and production judge defaults |
src/l2_bench_eval/task.py |
Builds the inspect-ai Task; fetches data from the Hub |
src/l2_bench_eval/dataset.py |
Parses l2-bench_tasks.csv + resources into a MemoryDataset |
src/l2_bench_eval/criteria.py |
Rubric parsing and per-task criterion lookup |
src/l2_bench_eval/score.py |
LLM-as-judge scorer and weighted aggregation |
src/l2_bench_eval/prompts/ |
Judge prompt templates (v1βv4) and retry suffix |
src/l2_bench_eval/bedrock_patch.py |
Raises the Bedrock client read timeout for long responses |
Licence
This reference implementation pipeline is released under the MIT Licence. Β© 2026 Oxford University Press.
The L2-Bench dataset is licensed separately under CC-BY-SA-4.0; see the dataset card.
Citation
If you use L2-Bench in your research, please cite our work. The full citation is available below:
@misc{edgell2026l2benchevaluationbenchmarkmeasuring,
title={L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education},
author={James Edgell and Wm. Matthew Kennedy and Ben Knight and Danielle Carvalho and Martin Ku and Isaac Pattis},
year={2026},
eprint={2607.08842},
archivePrefix={arXiv},
primaryClass={cs.CY},
url={https://arxiv.org/abs/2607.08842},
}
For any inquiries or feedback, including submitting corrections, please use the "Register Interest" webform on our L2-Bench site and we will respond as soon as we can.
Version History
| Version | Date | Author | Changes |
|---|---|---|---|
| 1.0.0 | 2026-07-30 | M. Ku | Initial open-source release |