You refined by

Full-time

Benchmark Jobs In Remote - 13,459 Job Positions Available

1 – 20 of 13,459 jobs
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  28 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  28 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  28 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  28 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  28 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  12 days ago
Automatically Get Matched to benchmark Jobs Let our AI agent search and match you to the best jobs from across the web
Try it now
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  12 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  12 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  12 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  12 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  12 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  12 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  12 days ago
LILT jobs

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt

LILT  12 days ago
ActiveFence jobs

Description You ship a benchmark every two to three weeks (example benchmark). Each one measures a frontier risk that nobody has measured yet. Some go public. Some go only to the labs. Some of the benchmarks and papers are

ActiveFence  11 days ago
Pathway jobs

About Pathway Pathway builds the first post-transformer frontier model that solves AIs fundamental memory problem. While transformers wake up in the same state every time—like Groundhog Day—our architecture enables true continuous learning, infinite context reasoning, and

Pathway  6 days ago
Weekday jobs

This role is for one of our clients Compensation: $44 - $56 per hour We are seeking experts in history and political science to author and review high-quality academic assessment content for an AI research initiative.

Weekday  5 days ago
Weekday jobs

This role is for one of our clients Compensation: $61 - $77 per hour We are seeking expert mathematicians to author and review high-quality academic assessment content for an AI research initiative. You will write and

Weekday  5 days ago
Turing jobs

Get notified about new Informatics Manager jobs in United States. 1,000+ Informatics Manager Jobs in United StatesManager of Clinical Research Data WarehousingTechnical Lead, Digital Health and AI IntegrationSenior Manager of Clinical Informatics - Hospital IT Department

Turing  17 days ago

Mercor is seeking an Applied Legal Benchmark Specialist for a remote contract role. You will author original law questions designed to test deep concepts and ensure clarity, then rate difficulty and provide correct answers with plausible distractors.

Visa Hunt  1 day ago

Subscribe for job alerts and resources to make your job search easier!

Confirmation email sent to

Check your email and click on the link to start receiving your job alerts

Receive the latest remote job openings for:

benchmark

You also might be interested in:

AI

Frontier

Datasets

Onboarding

Hybrid

Rubrics

Data Analysis

Benchmarking

API

Realistic

Confirmation email sent to

Check your email and click on the link to start receiving your job alerts

All Filters Apply
Sort by
Job Type
Employer/Recruiter