You refined by

Contract

Other Benchmark Jobs In Remote - 4,475 Job Positions Available

1 – 20 of 4,475 jobs
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  19 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  19 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  19 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  19 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  19 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  19 days ago
Automatically Get Matched to other benchmark Jobs Let our AI agent search and match you to the best jobs from across the web
Try it now
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  19 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  19 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  19 days ago
ActiveFence jobs

Description You ship a benchmark every two to three weeks (example benchmark). Each one measures a frontier risk that nobody has measured yet. Some go public. Some go only to the labs. Some of the benchmarks and papers are

ActiveFence  18 days ago
Pathway jobs

About Pathway Pathway builds the first post-transformer frontier model that solves AIs fundamental memory problem. While transformers wake up in the same state every time—like Groundhog Day—our architecture enables true continuous learning, infinite context reasoning, and

Pathway  13 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  3 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  3 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  3 days ago
LILT jobs
LILT ( Hong Kong )

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  3 days ago
WPP Media jobs

About WPP Media WPP is the trusted growth partner for the world’s leading brands. With exceptional talent, trusted data and intelligence, and world-class partnerships – all united by ourpioneeringagentic marketing platform, WPP Open – we help

WPP Media  28 days ago
Nebius jobs

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to

Nebius  28 days ago
YipitData jobs

About Us: YipitData is the leading market research and analytics firm for the disruptive economy and recently raised up to $475M from The Carlyle Group at a valuation over $1B. We analyze billions of alternative data

YipitData  28 days ago
GIC Private Limited jobs
GIC Private Limited ( Singapore )

GIC is one of the world’s largest sovereign wealth funds. With over 2,000 employees across 11 locations around the world, we invest in more than 40 countries globally across asset classes and businesses. Working at GIC

GIC Private Limited  28 days ago
Nium jobs

Nium provides global infrastructure for real-time cross-border payments. We were founded on the mission to deliver the global payments infrastructure of tomorrow, today. Our platform enables banks, fintechs, and global businesses to move money instantly, everywhere.

Nium  28 days ago

Subscribe for job alerts and resources to make your job search easier!

Confirmation email sent to

Check your email and click on the link to start receiving your job alerts

Receive the latest remote job openings for:

other benchmark

You also might be interested in:

AI

Hybrid

Equities

Onboarding

Fostering

Automation

Geographic

Finance

CRM

Wellness

Confirmation email sent to

Check your email and click on the link to start receiving your job alerts

All Filters Apply
Sort by
Job Type
Employer/Recruiter