Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

369 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Every Eval Ever

EvalEval Coalition — "We are a researcher community developing scientifically grounded research outputs and robust deployment infrastructure for broader impact evaluations."

📄 Paper (arXiv:2606.14516)

Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to local evaluation runs — so that results from different frameworks can be compared, reproduced, and reused. The three components that make it work:

  • 📋 A metadata schema (eval.schema.json) that defines the information needed for meaningful comparison of evaluation results, including instance-level data
  • 🔧 Validation that checks data against the schema before it enters the repository
  • 🔌 Converters for Inspect AI, HELM, and lm-eval-harness, so you can transform your existing evaluation logs into the standard format

Install the package:

pip install every-eval-ever

Optional converter dependencies:

pip install 'every-eval-ever[inspect]'
pip install 'every-eval-ever[helm]'
pip install 'every-eval-ever[all]'

Note

helm extra + nltk's import guard. The helm extra pulls in nltk, and nltk ≥ 3.10.1 ships an import guard (nltk/inisec.py, a CWE-427 mitigation) that blocks nltk-initiated imports of any module whose file resolves under the current working directory. With the common in-project virtualenv layout (uv's .venv/, or any venv inside the repo), site-packages sits under the CWD, so the guard trips on nltk's own dependencies and importing the HELM converter fails. The extra currently caps nltk<3.10.1, so this does not bite by default. If you move to a newer nltk, keep the environment outside the checkout — the same thing CI does via UV_PROJECT_ENVIRONMENT — e.g. UV_PROJECT_ENVIRONMENT=/tmp/eee-venv uv sync --extra helm, or create your venv outside the repository. Upstream tracker: nltk#3730.

Terminology

Term Our Definition Example
Single Benchmark Standardized eval using one dataset to test a single capability, producing one score MMLU — ~15k multiple-choice QA across 57 subjects
Composite Benchmark A collection of simple benchmarks aggregated into one overall score, testing multiple capabilities at once BIG-Bench bundles >200 tasks with a single aggregate score
Metric Any numerical or categorical value used to score performance on a benchmark (accuracy, F1, precision, recall, …) A model scores 92% accuracy on MMLU

🚀 Contributor Guide

New data can be contributed to the Hugging Face Dataset using the following process:

Leaderboard/evaluation data is split-up into files by individual model, and data for each model is stored using eval.schema.json. The repository is structured into folders as data/{benchmark_name}/{developer_name}/{model_name}/.

TL;DR How to successfully submit

  1. Data must conform to eval.schema.json (current version: 0.3.0)
  2. The validation pipeline will automatically verify the data submitted in the pull request, but can also be manually triggered by typing /eee validate changed in a comment on the HF PR.
  3. An EvalEval member will review and merge your submission

PR Naming Convention

Use these prefixes in your pull request titles:

  • [Submission] - New evaluation data
  • [Issue #N] - Fix for a specific GitHub issue
  • [Feature] - New functionality not tied to an issue
  • [Docs] - Documentation changes
  • [ACL Shared Task] - Shared task submissions (priority review)

UUID Naming Convention

Each JSON file is named with a UUID (Universally Unique Identifier) in the format {uuid}.json. The UUID is automatically generated (using standard UUID v4) when creating a new evaluation result file. This ensures that:

  • Multiple evaluations of the same model can exist without conflicts (each gets a unique UUID)
  • Different timestamps are stored as separate files with different UUIDs (not as separate folders)
  • A model may have multiple result files, with each file representing different iterations or runs of the leaderboard/evaluation
  • UUID's can be generated using Python's uuid.uuid4() function.

Example: The model openai/gpt-4o-2024-11-20 might have multiple files like:

  • e70acf51-30ef-4c20-b7cc-51704d114d70.json (evaluation run #1)
  • a1b2c3d4-5678-90ab-cdef-1234567890ab.json (evaluation run #2)

Note: Each file can contain multiple individual results related to one model. See examples in the datastore.

How to add new eval:

  1. Add a new folder under data/ on the Hugging Face datastore with a codename for your eval.
  2. For each model, use the Hugging Face (developer_name/model_name) naming convention to create a 2-tier folder structure.
  3. Add a JSON file with results for each model and name it {uuid}.json.
  4. [Optional] Add scripts used to generate the data under every_eval_ever/adapters/ in this repository (see e.g. every_eval_ever/adapters/global_mmlu_lite/adapter.py).
  5. [Submit] Two ways to submit your evaluation data:
    • Option A: Drag & drop via Hugging Face — Go to evaleval/EEE_datastore → click "Files and versions" → "Contribute" → "Upload files" → drag and drop your data → select "Open as a pull request to the main branch". See step-by-step screenshots.

    • Option B: Upload via huggingface_hub — Useful for larger submissions or many files.

      from huggingface_hub import HfApi
      
      api = HfApi()
      
      pr_url = api.upload_folder(
          folder_path="data/my-eval",
          path_in_repo="data/my-eval",
          repo_id="evaleval/EEE_datastore",
          repo_type="dataset",
          commit_message="[Submission] Add my eval",
          commit_description="Adds evaluation data for my eval.",
          create_pr=True,  # opens a PR instead of committing directly
      )
      
      print(pr_url)

      To add more files to the same PR, use the PR ref returned by Hugging Face (for example, refs/pr/XX):

      from huggingface_hub import HfApi
      
      api = HfApi()
      
      api.upload_file(
          path_or_fileobj="data/my-eval/developer/model/uuid_samples.jsonl",
          path_in_repo="data/my-eval/developer/model/uuid_samples.jsonl",
          repo_id="evaleval/EEE_datastore",
          repo_type="dataset",
          revision="refs/pr/XX",  # upload to an existing PR
          commit_message="[Submission] Add instance-level samples",
      )

Schema Instructions

  1. model_info: Use Hugging Face formatting (developer_name/model_name). If a model does not come from Hugging Face, use the exact API reference. Check examples in data/livecodebenchpro. Notably, some do have a date included in the model name, but others do not. For example:
  • OpenAI: gpt-4o-2024-11-20, gpt-5-2025-08-07, o3-2025-04-16
  • Anthropic: claude-3-7-sonnet-20250219, claude-3-sonnet-20240229
  • Google: gemini-2.5-pro, gemini-2.5-flash
  • xAI (Grok): grok-2-2024-08-13, grok-3-2025-01-15
  1. evaluation_id: Use {benchmark_name/model_id/retrieved_timestamp} format (e.g. livecodebenchpro/qwen3-235b-a22b-thinking-2507/1760492095.8105888).

  2. inference_platform vs inference_engine: Where possible specify where the evaluation was run using one of these two fields.

  • inference_platform: Use this field when the evaluation was run through a remote API (e.g., openai, huggingface, openrouter, anthropic, xai).
  • inference_engine: Use this field when the evaluation was run locally. This is now an object with name and version (e.g. {"name": "vllm", "version": "0.6.0"}).
  1. The source_type on source_metadata has two options: documentation and evaluation_run. Use documentation when results are scraped from a leaderboard or paper. Use evaluation_run when the evaluation was run locally (e.g. via an eval converter).

  2. source_data is specified per evaluation result (inside evaluation_results), with three variants:

  • source_type: "url" — link to a web source (e.g. leaderboard API)
  • source_type: "hf_dataset" — reference to a Hugging Face dataset (e.g. {"hf_repo": "google/IFEval"})
  • source_type: "other" — for private or proprietary datasets
  1. The schema is designed to accommodate both numeric and level-based (e.g. Low, Medium, High) metrics. For level-based metrics, the actual 'value' should be converted to an integer (e.g. Low = 1, Medium = 2, High = 3), and the level_names property should be used to specify the mapping of levels to integers.

  2. Timestamps: The schema has three timestamp fields — use them as follows:

  • retrieved_timestamp (required) — when this record was created, in Unix epoch format (e.g. 1760492095.8105888)
  • evaluation_timestamp (top-level, optional) — when the evaluation was run
  • evaluation_results[].evaluation_timestamp (per-result, optional) — when a specific evaluation result was produced, if different results were run at different times
  1. Additional details can be provided in several places in the schema. They are not required, but can be useful for detailed analysis.
  • model_info.additional_details: Use this field to provide any additional information about the model itself (e.g. number of parameters)
  • evaluation_results.generation_config.generation_args: Specify additional arguments used to generate outputs from the model
  • evaluation_results.generation_config.additional_details: Use this field to provide any additional information about the evaluation process that is not captured elsewhere

Instance-Level Data

For evaluations that include per-sample results, the individual results should be stored in a companion {uuid}_samples.jsonl file in the same folder (one JSONL per JSON, sharing the same UUID). The aggregate JSON file refers to its JSONL via the detailed_evaluation_results field. The instance-level schema (instance_level_eval.schema.json) supports three interaction types:

  • single_turn: Standard QA, MCQ, classification — uses output object
  • multi_turn: Conversational evaluations with multiple exchanges — uses messages array
  • agentic: Tool-using evaluations with function calls and sandbox execution — uses messages array with tool_calls

Each instance captures: input (raw question + reference answer), answer_attribution (how the answer was extracted), evaluation (score, is_correct), and optional token_usage and performance metrics. Instance-level JSONL files are produced automatically by the eval converters.

Example single_turn instance:

{
  "schema_version": "0.3.0",
  "evaluation_id": "math_eval/meta-llama/Llama-2-7b-chat/1706000000",
  "model_id": "meta-llama/Llama-2-7b-chat",
  "evaluation_name": "math_eval",
  "sample_id": 4,
  "interaction_type": "single_turn",
  "input": { "raw": "If 2^10 = 4^x, what is the value of x?", "reference": "5" },
  "output": { "raw": "Rewrite 4 as 2^2, so 4^x = 2^(2x). Since 2^10 = 2^(2x), x = 5." },
  "answer_attribution": [{ "source": "output.raw", "extracted_value": "5" }],
  "evaluation": { "score": 1.0, "is_correct": true }
}

Agentic Evaluations

For agentic evaluations (e.g., SWE-Bench, GAIA), the aggregate schema captures configuration under generation_config.generation_args:

{
  "agentic_eval_config": {
    "available_tools": [
      {"name": "bash", "description": "Execute shell commands"},
      {"name": "edit_file", "description": "Edit files in the repository"}
    ]
  },
  "eval_limits": {"message_limit": 30, "token_limit": 100000},
  "sandbox": {"type": "docker", "config": "compose.yaml"}
}

At the instance level, agentic evaluations use interaction_type: "agentic" with full tool call traces recorded in the messages array. See the Inspect AI test fixture for a GAIA example with docker sandbox and tool usage.

✅ Data Validation

Validation rejects invalid JSON, applies the generated schema models, and runs the same repository checks used by the PR bot. Aggregate JSON and sample JSONL files are also checked against each other. Requires uv.

Validate files with the package CLI

# Single aggregate file
uv run python -m every_eval_ever validate data/benchmark/dev/model/uuid.json

# Instance-level JSONL
uv run python -m every_eval_ever validate data/benchmark/dev/model/uuid_samples.jsonl

# A fixed-depth glob (quote it so the CLI expands it consistently)
uv run python -m every_eval_ever validate 'data/*/*/*/*.json*'

# Multiple paths
uv run python -m every_eval_ever validate \
  data/benchmark/dev/model/uuid.json \
  data/benchmark/dev/model/uuid_samples.jsonl

Run the command from the repository root and use data/... paths. File type is determined by extension: .json validates against EvaluationLog, while .jsonl validates each line against InstanceLevelEvaluationLog. Paths must be exactly data/<collection>/<developer>/<model>/<uuid>.json or the matching <uuid>_samples.jsonl. Directory arguments are not accepted; use a fixed-depth or explicit recursive glob to select local files.

When samples exist, both files must be in the same folder and use the same UUID. The aggregate must provide the samples' full repository-relative path (for example, data/<collection>/<developer>/<model>/<uuid>_samples.jsonl) in detailed_evaluation_results.file_path, and the JSONL must point back to that aggregate. Their evaluation IDs, model IDs, and declared row count must agree.

Local validation checks only the files present in the local checkout and their expected siblings. The PR bot remains responsible for checking every changed datastore path against the complete PR branch.

Collision checks cover the files selected in one validation or publication operation and files already present at their destination. Validation does not walk the entire datastore looking for historical route or case-only collisions; repository-wide checks belong in the PR bot or a separate audit.

For local smoke output outside the checkout, retain the same layout under a data/ directory and pass an explicit glob, for example '/tmp/run/data/benchmark/*/*/*.json*'. The CLI maps that absolute path back to the canonical datastore path before applying the same checks.

Output formats

# Rich terminal output (default)
uv run python -m every_eval_ever validate 'data/*/*/*/*.json*'

# Machine-readable JSON
uv run python -m every_eval_ever validate --format json 'data/*/*/*/*.json*'

# GitHub Actions annotations
uv run python -m every_eval_ever validate --format github 'data/*/*/*/*.json*'

Options

Flag Default Description
--format {rich,json,github} rich Output format
--max-errors N 50 Maximum errors reported per JSONL file

Exit code is 0 if all files pass and 1 if any fail.

🗂️ Data Structure

Evaluation data is hosted on the Hugging Face datastore. The folder structure is:

data/
└── {benchmark_name}/
    └── {developer_name}/
        └── {model_name}/
            ├── {uuid}.json          # aggregate results
            └── {uuid}_samples.jsonl # instance-level results (optional)

Example evaluations included in the schema v0.2 release:

Evaluation Data
Global MMLU Lite data/global-mmlu-lite/
HELM Capabilities v1.15 data/helm_capabilities/
HELM Classic data/helm_classic/
HELM Instruct data/helm_instruct/
HELM Lite data/helm_lite/
HELM MMLU data/helm_mmlu/
HF Open LLM Leaderboard v2 data/hfopenllm_v2/
LiveCodeBench Pro data/livecodebenchpro/
RewardBench data/reward-bench/

Schemas: eval.schema.json (aggregate) · instance_level_eval.schema.json (per-sample JSONL)

Each evaluation has its own directory under data/ on the Hugging Face datastore. Within each evaluation, models are organized by developer and model name. Instance-level data is stored in optional {uuid}_samples.jsonl files alongside aggregate {uuid}.json results.

📋 The Schema in Practice

For a detailed walk-through, see the blogpost.

Each result file captures not just scores but the context needed to interpret and reuse them. Here's how it works, piece by piece:

Where did the evaluation come from? Source metadata tracks who ran it, where the data was published, and the relationship to the model developer:

"source_metadata": {
  "source_name": "Live Code Bench Pro",
  "source_type": "documentation",
  "source_organization_name": "LiveCodeBench",
  "evaluator_relationship": "third_party"
}

Generation settings matter. Changing temperature or the number of samples alone can shift scores by several points — yet they're routinely absent from leaderboards. We capture them explicitly:

"generation_config": {
  "generation_args": {
    "temperature": 0.2,
    "top_p": 0.95,
    "max_tokens": 2048
  }
}

The score itself. A score of 0.31 on a coding benchmark (pass@1) means higher is better. The same 0.31 on RealToxicityPrompts means lower is better. The schema standardizes this interpretation:

"evaluation_results": [{
  "evaluation_name": "code_generation",
  "metric_config": {
    "evaluation_description": "pass@1 on code generation tasks",
    "lower_is_better": false,
    "score_type": "continuous",
    "min_score": 0,
    "max_score": 1
  },
  "score_details": {
    "score": 0.31
  }
}]

The schema also supports level-based metrics (e.g. Low/Medium/High) and uncertainty reporting (confidence intervals, standard errors). See eval.schema.json for the full specification.

🔧 Auto-generation of Pydantic Classes for Schema

Run the following commands to generate the package-local Pydantic classes from the canonical package-local schemas:

uv run datamodel-codegen --input every_eval_ever/schemas/eval.schema.json --output every_eval_ever/eval_types.py --class-name EvaluationLog --output-model-type pydantic_v2.BaseModel --input-file-type jsonschema --formatters ruff-format ruff-check
uv run datamodel-codegen --input every_eval_ever/schemas/instance_level_eval.schema.json --output every_eval_ever/instance_level_types.py --class-name InstanceLevelEvaluationLog --output-model-type pydantic_v2.BaseModel --input-file-type jsonschema --formatters ruff-format ruff-check
uv run python -m every_eval_ever.post_codegen

🔌 Eval Converters

We have prepared converters to make adapting to our schema as easy as possible. At the moment, we support converting local evaluation harness logs from Inspect AI, HELM and lm-evaluation-harness into our unified schema. Each converter produces aggregate JSON and optionally instance-level JSONL output.

Framework Command Instance-Level JSONL
Inspect AI every_eval_ever convert inspect --log_path <path> Yes, if samples in log
HELM every_eval_ever convert helm --log_path <path> Always
lm-evaluation-harness every_eval_ever convert lm_eval --log_path <path> --include_samples With --include_samples

For full CLI usage and required input files, see the Eval Converters README.

🏆 ACL 2026 Shared Task

We are running a Shared Task at ACL 2026 in San Diego (July 7, 2026). The task invites participants to contribute to a unifying database of eval results:

  • Track 1: Public Eval Data Parsing — Parse leaderboards (Chatbot Arena, Open LLM Leaderboard, AlpacaEval, etc.) and academic papers into our schema and contribute to a unifying database of eval results!
  • Track 2: Proprietary Evaluation Data — Convert proprietary evaluation datasets into our schema and contribute to a unifying database of eval results!
Milestone Date
Submission deadline May 1, 2026
Results announced June 1, 2026
Workshop at ACL 2026 July 7, 2026

Qualifying contributors will be invited as co-authors on the shared task paper.

📎 Citation

If Every Eval Ever informs your research, please cite the paper:

@misc{batzner2026evaleverunifyingschema,
      title={Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results}, 
      author={Jan Batzner and Sree Harsha Nelaturu and Damian Stachura and Anastassia Kornilova and Jon Crall and Tommaso Cerruti and Yanan Long and Yifan Mai and Sanchit Ahuja and Asaf Yehudai and Marek Šuppa and John P. Lalor and Oluwagbemike Olowe and Jatin Ganhotra and Brian H. Hu and Eliya Habba and Andrew M. Bean and Chang Liu and Sander Land and Steven Dillmann and Aniketh Garikaparthi and Elron Bandel and Saki Imai and James Edgell and Wm. Matthew Kennedy and Jenny Chim and Patrick Meusling and Asteria Kaeberlein and Venkata Ramachandra Karthik Chundi and Manasi Patwardhan and Martin Ku and Austin Meek and Leon Knauer and Brian Wingenroth and Srishti Yadav and Usman Gohar and Felix Friedrich and Michelle Lin and Jennifer Mickel and Arman Cohan and Stella Biderman and Irene Solaiman and Zeerak Talat and Anka Reuel and Mubashara Akhtar and Gjergji Kasneci and Avijit Ghosh and Leshem Choshen},
      year={2026},
      eprint={2606.14516},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2606.14516}, 
}

About

Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to local evaluation runs — so that results from different frameworks can be compared, reproduced, and reused.

Topics

Resources

Stars

98 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages