At Mitratech, we are a team of technocrats focused on building world-class products that simplify operations in the Legal, Risk, Compliance, and HR functions. We are a close-knit, globally dispersed team that thrives in an ecosystem
About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt
Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data DivisionReports to: CEO · Remote / LatAm / USThe role Evaluation is one of the hardest open problems in AI: we still dont have