Skip to content
@socialfoundations

Social Foundations of Computation

Max Planck Institute for Intelligent Systems, Tübingen

Popular repositories Loading

  1. whynot whynot Public

    A Python sandbox for decision making in dynamics

    Python 428 46

  2. folktables folktables Public

    Datasets derived from US census data

    Python 295 25

  3. tttlm tttlm Public

    Test-time-training on nearest neighbors for large language models

    Python 50 5

  4. lawma lawma Public

    Lawma: A lightly fine-tuned Llama model for legal classification tasks.

    Jupyter Notebook 34 1

  5. folktexts folktexts Public

    Evaluate uncertainty, calibration, accuracy, and fairness of LLMs on real-world survey data!

    Jupyter Notebook 30 6

  6. benchbench benchbench Public

    BenchBench is a Python package to evaluate multi-task benchmarks.

    Python 24 2

Repositories

Showing 10 of 18 repositories
  • whynot Public

    A Python sandbox for decision making in dynamics

    socialfoundations/whynot's past year of commit activity
    Python 428 MIT 46 6 0 Updated Aug 11, 2026
  • socialfoundations/benchmark-prediction's past year of commit activity
    Python 6 MIT 3 0 0 Updated Jul 16, 2026
  • folktexts Public

    Evaluate uncertainty, calibration, accuracy, and fairness of LLMs on real-world survey data!

    socialfoundations/folktexts's past year of commit activity
    Jupyter Notebook 30 MIT 6 0 1 Updated Jul 7, 2026
  • roc-n-reroll Public

    Code used for "ROC-n-reroll: How verifier imperfection affects test-time scaling" at ICLR 2026.

    socialfoundations/roc-n-reroll's past year of commit activity
    Jupyter Notebook 4 1 0 0 Updated Jul 7, 2026
  • socialfoundations/correct-looks-better's past year of commit activity
    Python 3 0 0 0 Updated May 29, 2026
  • mono-multi Public

    Code to reproduce the paper "Monoculture or Multiplicity: Which Is It?"

    socialfoundations/mono-multi's past year of commit activity
    Jupyter Notebook 1 MIT 0 0 0 Updated Oct 27, 2025
  • benchbench Public

    BenchBench is a Python package to evaluate multi-task benchmarks.

    socialfoundations/benchbench's past year of commit activity
    Python 24 MIT 2 1 0 Updated Oct 12, 2025
  • lm-harmony Public
    socialfoundations/lm-harmony's past year of commit activity
    Jupyter Notebook 6 MIT 0 0 0 Updated Sep 22, 2025
  • error-parity Public

    Achieve error-rate fairness between societal groups for any score-based classifier.

    socialfoundations/error-parity's past year of commit activity
    Python 19 MIT 3 0 2 Updated Aug 21, 2025
  • lm-evaluation-harness Public Forked from EleutherAI/lm-evaluation-harness

    A framework for few-shot evaluation of language models.

    socialfoundations/lm-evaluation-harness's past year of commit activity
    Python 1 MIT 3,583 0 0 Updated May 4, 2025

Top languages

Loading…

Most used topics

Loading…