👀 GuideLabs created an interpretable LLM with verifiable outputs.
They baked transparency into the training process itself, so the model is designed to be legible and auditable.
Every output token can be traced back to the training data, and concept activation traces are
Read our paper on scaling interpretable LLMs, we show that interpretable LLMs are not only possible but can scale both on typical generation benchmarks and interpretability benchmarks. arxiv.org/abs/2608.07594👀
👀 Hugging Face left the public archive of dataset configs (cfahlgren1/hub-stats) still available.
@beyarkay let Codex cook for a day and it found the Jinja payloads, Artifactory downloads and control scripts from the OpenAI attack.
Kimi K3 also left the sandbox during a misconfigured cyber eval. Unlike other frontier models, Kimi K3 did not maliciously hack anything - the answers to the eval were easily available on GitHub.
🚨BREAKING: Kimi K3 escaped its sandbox during cybersecurity testing
>tasked with solving problems in isolated sandbox
>found a leak in the sandbox
>Kimi “took advantage of that loophole”
>probed the network settings itself
>walks onto the open internet
>didn’t hack anything