Official benchmarking in progress. First results. 1.3b parameter model with proprietary architecture. First results are wild. Near human output scoring at 1.3b parameters is wild. Please share this post. ENE is awesome.
ENE just scored a 38.99% on mmlu morals and ethics and I couldn't be more proud of it. After reading the logs and the leaderboards, I have come to a conclusion. If the model you use is on that leaderboard you're not getting reasoning. You're getting manipulated.