We are already using assitant API to implement much better experience, and users love it, but we still have a big concern: We want to use GPT-4 turbo but we dont need 128k window for cost saving reason, but it seems there is no such limit option so it becomes really expensive.