Would it be crazy to think in the near future certain harness features could become learned model behaviours? Compaction would be a good candidate, or more broadly most of the clever "context engineering" tricks harnesses use the model could instead manage internally.
On harnesses, I vacillate between three beliefs:
- the less harness, the better. Models are the magic
- post training a model and harness is dramatically better and the model providers win
- harnesses have real independent value from the model
I have no idea which is right.
the 2nd wave of consumer ai is priesthood. i think people just want an oracle that tells them who they are, what to believe, how to look, who to love, if they're healthy, where they belong. that's why astrology, religion, looksmaxxing, wellness, identity will be the greatest
sudo sysctl iogpu.wired_limit_mb=X
where X = total_ram_mb - headroom_for_sys_mb
For headroom_for_sys_mb somewhere between 4-8GB is good, lower if you have nothing else on except the inference engine.
Without this, a 24GB M4 Pro is artificially limited to 16GB of VRAM.