MIT-licensed · Provider-agnostic · Multilingual

Evaluate open language models beyond a single score.

A reproducible toolkit for source-grounded Turkish generation, structured JSON, and translation consistency across English, German, Spanish, French, and Dutch.

Interactive preview

Benchmark explorer

Switch task and language to preview how the public toolkit presents quality, reliability and efficiency signals.

Sample data only. This static Space does not call a model API and does not present real benchmark results. It demonstrates the planned public evaluation interface without collecting prompts, credentials or user data.

Evaluation dimensions

Structured output preview


          
What it measures

Evidence, not vibes

The toolkit combines deterministic validators with operational metrics so model comparisons remain inspectable and repeatable.

✓

Schema compliance

Checks JSON validity, required fields and task-specific structures instead of relying only on free-form output review.

🔗

Source adherence

Verifies expected source URLs and highlights missing or malformed evidence links in source-grounded tasks.

🌍

Language consistency

Runs language heuristics across Turkish and five translation targets to identify mixed-language or incomplete outputs.

⏱

Latency and reliability

Records endpoint latency, failures, retries and response-format fallbacks for operational comparison.

▦

Portable reports

Exports machine-readable JSON and CSV alongside Markdown summaries suitable for reviews and public reports.

🧪

Safe mock mode

Includes deterministic synthetic fixtures so the evaluation workflow can be demonstrated without API keys or production data.

Public artifacts

Inspect and reproduce

The evaluation code is public. Future GPU-backed work is intended to add reproducible configurations, datasets and anonymized reports.