AfterQuery closes $30M Series A at $300M valuation

Read blog

Dataset catalog

LLM training datasets and open benchmarks

Browse AfterQuery's commercial training-data programs, custom dataset capabilities, and public research datasets for LLMs and AI agents.

Commercial availability and terms are confirmed per engagement. Public repositories link directly to their official source so you can review the current files, metadata, and license status.


Commercial LLM training datasets

These programs are named in AfterQuery's published research. Request current availability, scope, delivery format, and terms for your training objective.

Commercial · Request access

Office Agent Training Dataset

Professional-work tasks with file-grounded inputs, multi-step analysis, judged outputs, rubrics, and deliverables such as spreadsheets, documents, and reports.

Best for

SFT warmup and reinforcement learning for office, research, and professional-work agents.

Published evidence

NVIDIA used AfterQuery tasks for SFT warmup and pivot RL. Its published GDPval ablation rose from 35.3 without warmup to 46.7 with warmup.

Commercial · Request access

τ² customer-service and tool-use data

Multi-turn customer-service trajectories built around policies, APIs, tools, user simulation, and passing task outcomes across public and AfterQuery-created domains.

Best for

SFT and agent training for policy-following, tool selection, argument accuracy, and user coordination.

Published evidence

An AfterQuery experiment used 1,057 passing rollouts covering 500 unique tasks across six domain variants and reported gains of up to 4.33× on unseen official τ² test tasks.

Commercial · Request access

Terminal and coding-agent training data

Successful terminal-agent trajectories and verifier-backed RL tasks spanning exploration, planning, editing, testing, debugging, and task completion in isolated environments.

Best for

SFT and RLVR for terminal agents, SWE agents, tool use, debugging, and long-horizon coding workflows.

Published evidence

AfterQuery reported improving gpt-oss-20b from 3.1% to 17.0% on Terminal-Bench 2.0, with zero overlap between the training tasks and the official evaluation set.


Custom training-data programs

When an existing dataset does not match the target distribution, AfterQuery can scope a pilot around a defined capability, workflow, domain, rubric, or model failure.

  • Supervised fine-tuning demonstrations and reasoning traces

  • Preference, ranking, critique, and RLHF data

  • Rubric- and verifier-based reinforcement-learning tasks

  • Tool-calling, browser, and computer-use trajectories

  • Code, debugging, deep-research, and multimodal data

  • Custom evaluations and held-out benchmarks


How access works

  1. 01

    Share the model capability, workflow, domain, and evaluation target.

  2. 02

    Review a relevant commercial dataset or scope a custom pilot.

  3. 03

    Confirm delivery schema, training rights, security terms, quality criteria, timeline, and pricing for that engagement.

  4. 04

    Evaluate the pilot against the target capability before scaling production.


Open datasets and benchmarks

These repositories are publicly accessible. Check the linked source before production use: where no license is published, do not assume commercial reuse rights.

FinanceQA

Financial question-answering examples built from primary financial documents, with context, answers, reasoning, source links, and question types.

Published details
148 examples · CSV · English
License status
Apache License 2.0

App-Bench

Full-stack application-building tasks with product prompts and functional rubrics for evaluating coding agents and AI app builders.

Published details
6 application tasks · CSV · Prompts and rubrics
License status
No license published

UI-Bench prompts

Client-style briefs for evaluating the visual design quality of AI website builders, with a main set spanning websites and web apps across five categories.

Published details
30 main prompts · 20 websites · 10 web apps · 5 categories
License status
No license published

MCP-Universe

Finance and spreadsheet tasks with context files, golden outputs, domain labels, and the tools required to complete each workflow.

Published details
2 tasks · CSV · Finance and spreadsheets
License status
No license published

VADER

A human-evaluated benchmark for vulnerability assessment, detection, explanation, remediation, and test-plan generation in real software repositories.

Published details
174 vulnerability cases · 15+ languages · 75%+ multi-file
License status
CC BY 4.0

OmicsBench grader databases

A read-only SQLite bundle that supports deterministic, synonym-aware grading for computational-biology agent tasks.

Published details
23.6 GB repository · SQLite · Bioinformatics grading
License status
Repository metadata: MIT; source databases retain upstream licenses