“rumor-mills we’re following have been mentioning how the data industry is taking off in China — very much driven by American data companies selling to Chinese model labs. This could look like Chinese labs buying many of the same RL environments that are used by American frontier
GLM 5.3 notes and why we should stop being so surprised about these very strong Chinese models (most of this is talking myself through some of my denial -- yes, these models are the real deal).
interconnects.ai/p/glm-53-how-c…
for my final post as @waterloo_intern, i'd like to ‘outroduce’ myself.
it takes exactly 4 years and 8 months to make a waterloo intern.
today is the last day of mine.
the full life cycle, in four stages, is as follows:
1- interview hazing
2- cali or bust
3- inflection point
kimi paper readers in SHAMBLES after reading the deepseek paper (me, it’s me, and at least one more (henry, below))
as in, if you can get just as good (actually better of) a model with swa, how much of k3’s success can be attributed to its use of kda vs the remaining tricks, and
I (very embarrasingly) only realize sliding window attention also makes kv cache bounded (i.e. indep of seq len), similar to KDA
sure maybe it uses more kv cache than KDA but its not a magnitude more