You refined by

Temporary

Benchmark Jobs In Remote - 13,716 Job Positions Available

1 – 20 of 13,716 jobs
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  26 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  26 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  26 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  26 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  26 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  10 days ago
Automatically Get Matched to benchmark Jobs Let our AI agent search and match you to the best jobs from across the web
Try it now
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  10 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  10 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  10 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  10 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  10 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  10 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  10 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  10 days ago
ActiveFence jobs

Description You ship a benchmark every two to three weeks (example benchmark). Each one measures a frontier risk that nobody has measured yet. Some go public. Some go only to the labs. Some of the benchmarks and papers are

ActiveFence  9 days ago
Pathway jobs

About Pathway Pathway builds the first post-transformer frontier model that solves AIs fundamental memory problem. While transformers wake up in the same state every time—like Groundhog Day—our architecture enables true continuous learning, infinite context reasoning, and

Pathway  4 days ago
Weekday jobs

This role is for one of our clients Compensation: $44 - $56 per hour We are seeking experts in history and political science to author and review high-quality academic assessment content for an AI research initiative.

Weekday  3 days ago
Weekday jobs

This role is for one of our clients Compensation: $61 - $77 per hour We are seeking expert mathematicians to author and review high-quality academic assessment content for an AI research initiative. You will write and

Weekday  3 days ago
Turing jobs

Get notified about new Informatics Manager jobs in United States. 1,000+ Informatics Manager Jobs in United StatesManager of Clinical Research Data WarehousingTechnical Lead, Digital Health and AI IntegrationSenior Manager of Clinical Informatics - Hospital IT Department

Turing  15 days ago
OpenTeams jobs

OpenTeams is seeking a Senior AI/ML Test and Evaluation Engineer to build and operate benchmarking capabilities for AI platforms. You’ll develop automated metrics paired with human judgment, and document limitations and failure modes for senior stakeholders.

OpenTeams  1 day ago

Subscribe for job alerts and resources to make your job search easier!

Confirmation email sent to

Check your email and click on the link to start receiving your job alerts

Receive the latest remote job openings for:

benchmark

You also might be interested in:

AI

Hybrid

Datasets

Onboarding

Rubrics

Frontier

Data Analysis

Technical Ability

Microsoft Office

Agent

Confirmation email sent to

Check your email and click on the link to start receiving your job alerts

All Filters Apply
Sort by
Job Type
Employer/Recruiter