Methodology
Every number we publish follows the same standing rules. They exist so you don't have to trust us — you can check.
Evaluation rules
- Strong judges only. LLM-judged metrics use a frontier model at temperature 0, plus an independent second pass as a cross-check. We report judge agreement alongside every judged result.
- Structured scoring. Judges return typed, schema-validated output — never free-text impressions.
- Real tasks. Query sets come from real user logs or verified gold labels, not synthetic prompts written to be answerable.
- Full-scale confirmation. Improvements measured on samples are confirmed at full scale before we claim them.
- Honest metrics. We publish the runs that went against us (see the landmark result in our first finding). Judge-generous calls are counted for the other side, not ours.
- Permanent holdouts. Benchmark sets used as gates are frozen; nothing in the production stack trains on them.
Reproducibility
- Every finding links a repository with the harness: runner scripts, judge prompts, metric code, and the evaluation data needed to re-run it.
- German court decisions and statutes are public domain (§ 5 UrhG) — our corpus can be rebuilt from official sources. That makes our benchmarks fully reproducible in a way benchmarks on licensed case-law collections cannot be.
- Licensing: harness code under PolyForm Noncommercial 1.0.0, our questions and annotations under CC BY-NC 4.0. Commercial licensing on request.
Corpus and extraction
- Sources. The corpus behind gesetze.nulegal.eu aggregates officially published decisions and statutes: the federal courts' portal (Rechtsprechung im Internet), state justice portals (with NRW the largest), and EU sources. Statutes come from the federal publication pipeline with full version history.
- Coverage is uneven by design of the sources. The federal portal is essentially complete from about 2010; earlier years and state coverage vary widely. Counts on the data pages measure published decisions in this corpus, never total judicial output.
- Citation extraction parses citations from decision full text (case-to-case and case-to-statute). It is rule-based; we have not yet published a formal precision/recall audit, so treat exact edge counts as close approximations rather than ground truth. An audit is on the roadmap.
- Attribution gaps are disclosed. Decisions whose court cannot be attributed are excluded from per-court statistics and reported as an explicit remainder in each snapshot.
Editorial rules
- No comparisons naming competing legal-research products. We benchmark against general-purpose baselines anyone can access (for example Google's own search products).
- Curated paper pages carry our own written summaries, never republished abstracts.
- Corrections are made in place and dated, not silently.
- Drafting is AI-assisted; every published claim is human-reviewed and every number comes from a logged run.