Very exciting paper! Eventually misalignment is a property of improper training and we need to figure out how to train better: guardrails are a useful posthoc patch and will be part of a broader safety ecosystem, but the root problem here is training.
Can training against lie detectors make frontier LLMs more honest? Our latest paper says yes. As models scale from 1B to 405B params, undetected deception more than halves. 1/8




