0

PatentQATrain

Fresh

PatentQATrain is a training environment for patent question answering, based on the PatentQA task from LAB-Bench-2. Agents are given questions about specific details from patents across diverse technology domains and must use web search to find and verify answers from Google P…

Type
RL Env
Runtime
ORS
License
unknown
Size
796 tasks
Published
Mar 2026

Cite

Notes

Only stored in your browser.

PatentQATrain

⭐ OpenReward Environment

Description

PatentQATrain is an ORS training environment for patent question answering, based on the PatentQA task from LAB-Bench-2. Agents are given questions about specific details from patents across diverse technology domains and must use web search to find and verify answers from Google Patents.

Capabilities

  • Researching patent details using web search
  • Extracting specific information from patent claims and specifications
  • Navigating patent documents to find numerical thresholds, compositions, method steps, and claim elements
  • Cross-domain patent understanding spanning pharmaceutical, biotech, chemistry, electronics, software, mechanical, materials, energy, medical devices, and telecom

Compute Requirements

No sandbox or special compute requirements. Uses external web search for patent retrieval.

License

MIT

Tasks

There are 909 training tasks distributed across 10 patent domains:

DomainCountDescription
pharmaceutical90Drug compositions, formulations, drug delivery
biotech91Genetic engineering, antibodies, cell therapy
chemistry89Chemical compounds, synthesis, catalysis
electronics97Semiconductors, circuits, displays, sensors
software92Algorithms, data processing, networking
mechanical95Engines, mechanisms, manufacturing
materials86Polymers, composites, coatings, nanomaterials
energy88Solar cells, batteries, fuel cells
medical_devices89Surgical instruments, imaging, prosthetics
telecom92Wireless protocols, signal processing

Every task's key_passage is verbatim text from the patent at its source_url, and every answer was checked against that patent by an LLM judge (99.6% supported). Patents where no passage produced a verifiable question were dropped, which is why the counts sit below 100.

Expect some drift. Both discovery and reading go through a live web-search API, and Tavily's extraction of a Google Patents page varies between calls — the same URL has returned 12k and 132k characters minutes apart. So a task that is answerable today can become unanswerable if extraction stops returning the section its answer came from, independently of anything in this repo. Verification here was done against extracts captured at generation time, which is the optimistic side of that: a row proven verifiable then is not guaranteed reachable now.

Each task presents a question about a specific patent that requires distinctive specificity — it asks about a verifiable fact uniquely traceable to one patent. Roughly 15% of questions cite the patent number outright; the rest describe the invention and leave the agent to identify the patent from that description, which is the harder and more realistic research task.

Reward Structure

Sparse, binary reward:

  • 1.0 for correct answers (as judged by LLM grader)
  • 0.0 for incorrect or unsure answers

Grading uses semantic equivalence checking: answers that are numerically/semantically equivalent are accepted, even if phrased differently. The grader is based on LABBench2's structured evaluation prompt.

We do not use exact string matching. The LLM grader (gpt-5-mini) evaluates whether the submitted answer captures the core factual content of the expected answer.

Data

Ground-truth data consists of QA pairs derived from Google Patents documents. Each task includes:

  • A question about one specific patent, sometimes citing its number and otherwise describing the invention
  • A concise expected answer (under 200 characters)
  • The source patent URL
  • A key passage from the patent supporting the answer

Data is stored on the OpenReward platform.

Tools

Agents have access to three tools:

ToolDescription
web_searchSearch the web. Returns titles, URLs, and snippets.
web_fetchFetch full text content from a URL via the configured search backend. Long documents are paginated (~25,000 chars/page), and search="<keyword>" returns the matching excerpts with their page numbers so the agent can jump to the relevant part instead of paging through. Each document is extracted once per session and cached.

Grading runs through a hidden @terminal tool rather than a tool the agent can call: replying with a plain message ends the rollout, and that message text is graded against the expected answer.

Time Horizon

PatentQATrain is a multi-turn environment. Agents typically perform several web searches and URL fetches before submitting an answer.

Environment Difficulty

[Statistics on environment difficulty here]

Other Environment Requirements

This environment requires the following API keys passed via the secrets parameter:

  • openai_api_key: For LLM-based answer grading
  • Search credentials — whichever the configured backend needs: api_key for the default backsearch backend, or tavily_api_key when the server runs with OPENREWARD_SEARCH_BACKEND=tavily. Both fall back to the server process environment (OPENREWARD_API_KEY / TAVILY_API_KEY).

Safety

PatentQATrain focuses on factual information retrieval from publicly available patent records. The environment does not involve intellectual property creation, legal advice, or access to non-public information. All source patents are publicly accessible via Google Patents.

Citations

This environment is inspired by the PatentQA task from LAB-Bench-2:

@misc{labbench2,
  author    = {Laurent, Jon M. and Bou, Albert and Pieler, Michael and Igoe, Conor and Andonian, Alex and Narayanan, Siddharth and Braza, James and Vassopoulos, Alexandros Sanchez and Steenwyk, Jacob L. and Lash, Blake and White, Andrew D. and Rodriques, Samuel G.},
  title     = {LABBench2: An Improved Benchmark for AI Systems Performing Biology Research},
  year      = {2026},
  url       = {https://github.com/EdisonScientific/labbench2}
}