Skip to main content

Quick Tips

Start Small

Always test with --limit 10 before running full benchmarks

Use Model & Task Flags

Use -M for model and -T for any benchmark-specific arguments

Debug Mode

Use --debug for full stack tracing when troubleshooting

Detailed Breakdown

Use bench view for detailed sample-by-sample evaluation breakdown

Global Help

Use --help on any command to see all available options

Use Groq for Testing

Free tier with fast inference - perfect for development

Common Issues & Solutions

Package not properly installed, try:
Remember: Command-line arguments override environment variables
The reasoning_effort parameter is now a first-class CLI flag.

Runtime Errors

Still Need Help?

GitHub Issues: Report bugs or ask questions