Optimizing between the harness and inference layer is an exciting domain that is woefully under explored to provide over the top token and speed efficiencies without sacrificing output results.
Much more to do here, but excited to see this trend.
Replying to @pwendell and @databricks
Full post here:
databricks.com/blog/managing-β¦




