AI & ML
How We Cut Target Leakage from 92% to 0.1% in Court Outcome Prediction (and Built a Triple-A MCP Server)
Elio Coppola DEV Community
4 views
When building AI for court outcome prediction, there is a massive hidden trap that invalidates most benchmarks: target leakage.
In Dutch court rulings, roughly 92% of raw texts contain the actual outcome verbatim (dictum or conclusion sentences like "the court dismisses the claim"). If you feed raw text to a model, it does not learn legal logic. It simply learns to read the answer back to you.
Here is how we solved this across 609,715 cases, built a traceable LightGBM model, and exposed it as a zero-dependency open-core MCP server.
1. The Pre-Training Cut: 92% to 0.1% Leakage
Before training any classifier, we implemented a strict sanitization step:
The dictum, summary lines, and outcome-announcing phrases are stripped from the text.
We continuously measure residual outcome markers.
Result: Leakage dropped from 92% to 0.1% (around 1 in 1,000 texts).
Only on this sanitized dataset did we train.
2. Why LightGBM Instead of an LLM
We intentionally picked LightGBM over deep neural networks or fine-tuned LLMs:
Fast and cheap: Sub-10ms inference without GPUs.
Traceable: Clear tree structures and feature importance.
Deterministic calibration: If confidence drops below 55%, the model does not guess. It returns "insufficient certainty".
3. Benchmark on 609,715 Cases (Out-of-Fold)
Evaluated through 5-fold cross-validation, strictly measured out-of-fold:
Overall Accuracy: 78.2% (against a 43.7% majority baseline)
Macro-F1: 77.1%
Per-Class F1:
Dismissed: 0.827
Partly granted: 0.726
Granted: 0.761
Domain Breakdown:
Criminal Law (n=105,151): 82.6% accuracy, 0.804 macro-F1 (strongest performance)
Administrative Law (n=316,273): 81.0% accuracy, 0.685 macro-F1 (high accuracy, but government victory is the majority class)
Civil Law (n=188,177): 71.0% accuracy, 0.656 macro-F1 (most complex due to factual nuances)
4. Model Context Protocol (MCP) Interface
To make this accessible to AI assistants (Claude, Cursor, autonomous agents), we wrapped the pipeline into an MCP server.
Zero external dependencies: Single Python file using only standard library (sys, json, urllib).
Audited on Glama: Triple-A rating (5/5 on coherence and completeness).
3 Tools:
rechtspraak_cijfers (keyless): Benchmark statistics and baseline metrics.
lekkage_check (keyless): Paste any legal text to test for outcome leakage before and after the cut.
voorspel_uitkomst (key required): Outcome risk classification.
Links
GitHub Repo (Apache 2.0): https://github.com/rechtssysteem-ai/rechtssysteem-mcp
Live Benchmark: https://rechtssysteem.ai/benchmark
Glama Audit: https://glama.ai/mcp/servers/@rechtssysteem-ai/rechtssysteem-mcp
Keyless API: https://api.rechtssysteem.ai/cijfers
Disclaimer: Not legal advice. Built as an open, verifiable yardstick for legal tech developers and agents.
Read original: https://dev.to/rechtssysteem/how-we-cut-target-leakage-from-92-to-01-in-court-outcome-prediction-and-built-a-triple-a-mcp-478g
← Previous
Why Client-Side Tracking Fails in Fintech (And How to Implement Meta CAPI for Funded Accounts)
Next →
Upgrade Skill Development Kamu dengan Menjelajahi Fitur Keren di Tencent EdgeOne Makers
Related
Agent Guardrails Beat Agent Capability: Three September Incidents Every Cross-Border Seller Should Read
AI & ML
5
DEV Community
Free Rails Architecture Checkup: run it in Codex, Claude Code, Cursor, etc. It shows where your app gives agents conflicting instructions, hides key architecture, or leaves risky decisions up to guesswork. https://railsbaseline.com/checkup/try/
AI & ML
3
Dev.to (EN Zone)
Quantified Self: Transform Your Medical PDFs into a Personal Health Oracle with RAG & PubMed
AI & ML
5
Dev.to (EN Zone)
Beyond the Hype: How ‘AI Psychosis’ and the OpenAI Agents API Are Exposing the Fragility of Modern Software Engineering
AI & ML
4
Dev.to (EN Zone)
Comments0
No comments yet — be the first