I like this idea a lot. Simple and clean. Rather than just having an equal weight residual sum of transformer layers, let each layer use attn to compute an weight for each previous layers residual. Allowing each layer of the transformer to build an appropriate input by weighting
- Cohere-transcribe with a whisperX style interface: github.com/Diffio-AI/Cohe…
- Codex Wrapped 2025 Total Tokens: 3,073,600,806 Total Messages: 1,782 Total Sessions: 512 Top model: GPT 5.2 Codex Total Estimated Cost: $814.14 Credit: @nummanali @moddi3io

