About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt
Build the GenAI platform that powers critical decisions in healthcare, legal, tax, and compliance industries. Your work will directly shape the future of these fields, enabling faster, safer, and more impactful decision-making at a global scale. -Location: