OdenseNLP

OdenseNLP

Safe, Efficient and Open Natural Language Processing @ University of Southern Denmark

Latest news

See all posts
DFM Mimir v1 Release

14 August 2026

DFM Mimir v1 Release

Danish Foundation Models has released Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture and trained from scratch on permissi...

OdenseNLP FlexMoRE image

02 July 2026

FlexMoRE - Efficient modular language models

The Allen Institute for AI (Ai2) recently highlighted FlexMoRE, a new approach to building more efficient modular language models developed by researchers at OdenseNLP and collaborat...

OdenseNLP LREC26 image

16 May 2026

OdenseNLP at LREC26

OdenseNLP was at LREC 2026 and authored/contributed in three papers:

Current focuses

Low-resource NLP

Language technologies for low-resource langauges, particularly Danish and neighboring Scandinavian languages.

Efficient NLP

Fast and efficient NLP architectures and methods.

AI Safety & Interpretability

Making AI systems more safe, trustworthy, and interpretable.

Datasets

Post-Training Dataset Collection

An extensive collection for language-model training with objectives spanning English and Danish instruction and knowledge, mathematics, and agentic...

Danish Dynaword

A continually expanded corpus of openly licensed Danish free-form text from diverse domains. Version 1.2.20 contains 7.25 million documents and 6.9...

Benchmarks

DaLA

Danish linguistic acceptability benchmark with corrupted and non-corrupted sentences, published as ~8.68k examples with train/validation/test and f...

SDU-Daisy

Danish-culture benchmark based on the Danish Culture Canon, with 746 closed question-answer pairs for evaluating LLM cultural understanding.

GEC DaLA

Grammatical error correction version of DaLA, pairing original and corrupted Danish sentences with corruption types and affected-token annotations....

DFM Mimir

An instruction-tuned Danish-English foundation model based on HRM-Text, trained exclusively on permissible post-training data and designed for stro...

Danish Foundation Models

The complete, growing collection of Danish Foundation Models releases on Hugging Face, spanning text-generation and multimodal models, Danish-focus...

Qwen3.5-27B-psysafe

A supervised fine-tune of Qwen3.5-27B for psychologically informed refusals, providing structured, supportive responses to high-risk requests while...

Repositories

See all repositories

PsychoSafe

Pipeline for building and evaluating psychologically informed refusal behavior in large language models through data creation, prompting, fine-tuni...

PropMe

Framework for evaluating memorization and propensity-aware memorization of training data in large language models.

DaLA

Source code for creating the Danish Corpus of Linguistic Acceptability, designed to evaluate Danish linguistic acceptability with real-world errors.