🌏 中文版
This post is written from the slides; I will add to it once the video is posted. As of 2026-09-29, CMU 11-768 has released only the 124-page slide deck for Lecture 10, with no recording. Every number and example below comes from the slides or the papers they cite. For each paper used to support a point, I opened the full text and checked the relevant passage: where the slides and the paper differ, both are given, and anything found only on the slides is marked as such. What the lecturer said out loud in class is unknown until the video is out, and I do not guess at it here.
CMU 11-768 AI Agents is Daniel Fried and Graham Neubig's Fall 2026 course on agents. Its Domains module covers coding agents, then computer use agents, and the third domain is deep research. Lecture 10 (Sep 24) is a guest lecture by Akari Asai, first author of OpenScholar (arXiv:2411.14199, Nature 2026) and joint first author of DR Tulu (arXiv:2511.19399, ICML 2026), so half of this lecture is her explaining how she built these systems.
A deep research agent takes a research question that needs many searches and synthesis across many documents, then plans, searches, reflects, searches again, and finally hands back a long answer with citations. The slides split the lecture into three parts: evaluation (benchmarks, rubrics, citation support), modeling (learning to search and synthesize), and retrieval (finding evidence for the agent's next step). This post follows the same order.
One search versus many
The lecture opens with a contrast. "What is Akari Asai's office number at CMU?" takes one search of the faculty page. "Can AI agents synthesize scientific literature as well as human experts?" does not, and the slides break the agent's work on it into four steps:
- Plan: separate "strong benchmark results" from "direct comparisons with human experts," then plan to find literature-synthesis studies, check how experts and agents were compared, and compare tasks and limits.
- Search: two kinds of results come back — a benchmark that scores answers against other AI systems, and an expert evaluation that compares agent answers with human-written ones.
- Reflect: the two studies measure different things, and a higher benchmark score does not answer the question about human experts. So the agent asks which tasks and domains were tested, and searches again.
- Final answer: "Promising on tested tasks; broader parity with experts remains unproven," citing OpenScholar and DR Tulu separately.
Step three is the point. The hard part of deep research is not searching many times. It is judging whether the evidence supports the conclusion, then narrowing the conclusion to what the evidence can hold. The rest of the lecture's evaluation material is about measuring exactly that.
Evaluation: four gaps between QA and deep research benchmarks
The slides use a Natural Questions item as the foil: "who wrote the score for the force awakens?" The answer is John Williams, found in one Wikipedia page. Classic QA benchmarks like this have four gaps, and the section fills them one at a time:
| Gap | Problem with classic QA | Benchmarks the slides use to fill it |
|---|---|---|
| Search complexity | One search, or recall, is enough | BrowseComp, BrowseComp-Plus |
| Domain expertise | Broad web questions do not test expert knowledge | MedBrowseComp, FinSearchComp |
| Answer quality | Short-answer matching cannot assess multi-document synthesis | ScholarQABench, ResearchQA |
| Evidence support | A correct answer does not show the evidence was used correctly | DeepResearch Bench |
Search complexity: BrowseComp writes questions backwards
BrowseComp (arXiv:2504.12516) has 1,266 human-written questions. The authors write them backwards: start from a known person, event, or artifact as the answer, hide its name, and combine several factual clues into one complex question. Difficulty is checked twice: models with search should fail and five Google searches should not suffice, and a human should still find it hard after 10 minutes. The paper notes the 10-minute rule was not strictly enforced; a second trainer attempted only a portion of the questions.
The slides' example asks for a fictional character. The question gives five clues, covering the character's narrative device, backstory, personality, and when their TV show aired and how many episodes it ran. Each clue on its own returns a pile of candidates; only the intersection converges. (The BrowseComp paper asks readers not to repost its example questions in plain text, so they don't leak into training data and contaminate the benchmark. This post therefore describes only the question's structure, not the question or its answer.)
The results on the slide (the five model rows match Table 3 of the BrowseComp paper):
| System | Accuracy |
|---|---|
| GPT-4o | 0.6% |
| GPT-4o + browsing | 1.9% |
| GPT-4.5 | 0.9% |
| OpenAI o1 (medium) | 9.9% |
| Deep Research (trained for this kind of task) | 51.5% |
| Human reference match (computed on the slide) | 25.3% |
The paper does not report the human row this way. Its numbers: trainers attempted 1,255 questions and solved 29.2%, giving up on the other 70.8% after two hours of searching; of the solved questions, 86.4% matched the reference answer. The slide's 25.3% multiplies the two (317 / 1,255).
Adding browsing barely helps; the jump comes from training the search behavior itself. The slides then show Figure 10(a) of the Tongyi DeepResearch report (arXiv:2510.24701): as context length grows from 8K to 128K, BrowseComp accuracy climbs from near zero to over 40% (values read off the plot). The report also says over 20% of its SFT samples exceed 32K tokens and involve more than 10 tool calls, and evaluation allows up to 128 tool calls per task. Give the agent more room to search and accuracy follows.
BrowseComp runs on the live web, where search APIs and pages keep changing, so two systems' scores are hard to compare fairly. BrowseComp-Plus (arXiv:2508.06600) freezes the corpus. It takes BrowseComp question–answer pairs, has o3 find evidence pages, has humans mark the spans that support each clue and the answer documents, then adds hard negatives. The slides describe them as related pages that miss a clue; in the paper, GPT-4o splits each question into about seven sub-queries, each is sent to Google search, and the returned pages become distractors. Of the original 1,266 questions, 830 passed human verification, and the fixed corpus is 830 questions over 100,195 documents. Once the corpus is fixed, you can isolate the retriever's contribution, which matters in the last section.
Domain expertise: let experts write the questions
Broad web questions do not test expert judgment. In FinSearchComp (arXiv:2509.13160; ICLR 2026 per the slides), finance experts write questions from work scenarios or financial tables and cross-check answers against multiple sources; one or two other experts then solve each question blind, and a senior expert arbitrates when results disagree. The paraphrased example: how did Johnson & Johnson's international share of revenue change year over year during 2022–2024? The same idea appears in medicine with MedBrowseComp (linking trials, drugs, and regulatory facts), and in academic search with ScholarSearch and AutoResearchBench (find one paper, or find all matching papers).
Answer quality: similarity metrics fail, so use rubrics
The slides reuse the opening question to show why long answers are hard to grade. A reference summary says "experts preferred OpenScholar in one study; parity across domains is unproven." Another valid summary says "one study favors OpenScholar; it does not establish equivalence across domains." A third says "…parity across domains is proven" — nearly identical to the reference in wording, and wrong. There are several valid answers, and similarity to the gold answer is not enough.
The evidence agrees. The slides cite A Critical Evaluation of Evaluations for Long-form Question Answering (ACL 2023): across 109 expert comparisons with reference answers, using a similarity metric to pick the answer experts preferred gets ROUGE to 58%, BERTScore to 57%, and BLEURT to 62%, against a 50% chance baseline. The same table has a more embarrassing comparison: on the same expert comparisons (129 pairs, including fields with no reference answer), simply picking the longer answer scores 68%, higher than all three metrics.
The alternative is a rubric. In ScholarQABench, from the OpenScholar paper, the Scholar-CS subset (100 questions) has PhD-level experts list the ingredients a good answer needs, marked must-have or nice-to-have; GPT-4o checks each one, and this part is 60% of the score, with the other 40% covering general criteria such as length, expertise, citations, and excerpts. The slides simplify this into a two-item weighted example: "reports a direct comparison with human experts" (weight 2) and "qualifies the conclusion by task and domain" (weight 1). A candidate that says "OpenScholar wins 70% of expert comparisons, and this holds across all fields" passes the first and fails the second, scoring (2×1 + 1×0) / 3 = 0.67. The Nature version of the paper measures agreement on whether a rubric item is satisfied, using outputs from two systems and two expert annotators: 0.80 between the experts and 0.79 between an expert and the LLM judge. The slides add that this covers 12 questions with 2 expert raters per answer; the question count is not in the paper's main text, so it follows the slides here.
Expert-written rubrics do not scale. ResearchQA (arXiv:2509.00496; TACL 2026 per the slides) mines questions from survey papers and uses LMs to generate rubrics, reaching 21K questions across 75 fields with 160K rubric items. The cost is that the rubric itself can be wrong. In the expert audit, an item is void if it is hard to judge, unclear, vacuous (for example, a mere rephrasing of the query), or erroneous (for example, citing a nonexistent paper). Of the three rubric types, parametric rubrics — generated from the model's own knowledge — had the highest void rate, 15%; the paper's final benchmark uses hybrid rubrics that merge survey content with model knowledge. The slides add one more failure: a vague criterion lets plausible errors pass. The slides' takeaway: evaluate the rubric before using it to evaluate an answer — ground criteria in sources, then audit their relevance and verifiability. DeepResearch Bench II does this: it derives criteria from expert-written investigative reports through four stages (LLM extraction, LLM self-evaluation, manual revision, and domain-expert review), ending with 132 tasks and 9,430 binary rubric items.
Evidence support: check every citation
An answer can be right and still cite the wrong things. The FACT framework in DeepResearch Bench (arXiv:2506.11763; ICLR 2026 per the slides) has three steps: an LLM extracts each statement with its cited URL and merges duplicates, the cited page's text is fetched, and an LLM judges whether the page supports the statement; citation accuracy and average effective citations per task are computed from those judgments. In the slides' own example, "OpenScholar was compared with human answers [1]" is supported by the OpenScholar paper; "OpenScholar matches experts in every field [1]" is not, because the paper did not test every field. One of two claims holds, so citation accuracy is 50% and the effective citation count is 1.
With all four gaps filled, the slides' summary grid reads: the BrowseComp family covers search difficulty, MedBrowseComp and FinSearchComp cover domains, ScholarQABench and ResearchQA cover answer quality with rubrics, and DeepResearch Bench covers citation support. No single benchmark covers all four cells, so before picking one, decide which cell you are trying to measure.
Modeling: how a deep research agent is trained
The inference loop
The slides first unpack one inference: the model generates <think> reasoning → generates a <tool>search(...)</tool> call → the search runs → retrieved text is appended to the context → the model continues reasoning with the new evidence → it produces an <answer> with citations. This is the tool-calling loop from L2; the only differences are that the tool is search and the output is a cited long answer.
A three-stage recipe: mid-training, SFT, RL
The slides use Tongyi DeepResearch's pipeline as the template:
| Stage | Goal | Training data |
|---|---|---|
| Mid-training | Learn agent behavior at scale | Input + answer, teacher trajectories |
| SFT | Imitate successful demonstrations | Input + answer, teacher trajectories |
| RL | Learn from feedback on its own attempts | Input + answer, the policy's own rollouts |
The next question is where the tasks and trajectories come from.
Synthesizing tasks by sampling a knowledge graph. The slides cite WebSailor-V2 (arXiv:2509.13305), which builds a dense knowledge graph, extracts subgraphs by random walk, and writes questions from them. The CMU example below is the slides' own illustration, with facts from the CMU archives and the Nobel Prize records. Nodes are entities and labeled arrows are relationships: Allen Newell was CMU faculty; Newell and Herbert Simon coauthored Human Problem Solving (1972); they shared the 1975 Turing Award; Simon won the 1978 economics Nobel Prize. Sample a connected subgraph, hide the target's name, and turn the relations into clues: "Which CMU researcher coauthored a book on human problem solving with a future economics Nobel laureate, and shared the Turing Award with that colleague three years before the Nobel Prize?" The answer, Allen Newell, is known in advance, so it can be checked automatically. It is the same logic as BrowseComp's hand-written questions, produced by machine at scale.
SFT: collect, filter, then imitate. A teacher model generates trajectories for these questions, and only those with a correct answer, a valid format, and a coherent trajectory are kept for SFT. This matches WebDancer's three-stage filter (validity, correctness, quality); the Tongyi report calls it rejection sampling. Correct-but-malformed and well-formed-but-wrong trajectories are both dropped.
Mid-training: get the base model used to agent workflows first. Tongyi's mid-training is Agentic Continual Pre-training (Scaling Agents via Continual Pre-training; ICLR 2026 per the slides): continue pretraining the base model on large amounts of synthesized agent-behavior data plus web text, tool-call records, and discarded trajectories. The slides compare two starting points with the same downstream SFT data (Pass@1, %; the paper's Table 3 SFT-B setting, with Qwen3-30B-A3B and AgentFounder-30B):
| Benchmark | Qwen3 base + SFT | AgentFounder base + SFT |
|---|---|---|
| BrowseComp-en | 28.6 | 39.9 |
| BrowseComp-zh | 35.6 | 43.3 |
| GAIA | 71.8 | 72.8 |
A better starting point carries through SFT. The paper tries three SFT datasets, and the BrowseComp-en gap is 4.5 (SFT-A), 11.3 (SFT-B), and 14.3 (SFT-C) points; the slides show the middle one.
RLVR: answer correctness as the reward
For short-answer questions, correctness can be the reward directly. Search-R1 (arXiv:2503.09516) uses only final-answer exact match as its reward. The slides' illustration (the question and rewards are the slides' own): the policy searches and reasons on one question, three sampled trajectories answer Herbert Simon, Allen Newell, and Allen Newell, matching against the known answer gives rewards 0, 1, 1, and those rewards update the policy.
The slides' update rule is GRPO from DeepSeekMath (the Search-R1 paper tries both PPO and GRPO): compare attempts within a group for the same question, raise the probability of those above the group average and lower those below. L9 covered the basic RL objective (maximize expected reward); GRPO in detail is the topic of L11.
The GRPO objective (as shown on the slides)
$$ J(\theta) = \mathbb{E}\left[\min\left(\rho \hat{A},\ \mathrm{clip}(\rho, 1-\epsilon, 1+\epsilon)\hat{A}\right) - \beta D_{\mathrm{KL}}\right] $$
- $\hat{A}$: relative reward within the group (compared with other attempts at the same question)
- $\mathrm{clip}$: limits the size of each update
- $\beta D_{\mathrm{KL}}$: keeps the policy near the reference model
The slides note the loss is computed only on model-generated tokens, not on tool output. This is Search-R1's retrieved-token loss masking.
Asynchronous rollouts. Trajectories use different numbers of tool calls; some finish in two turns, others take five. The slides score each trajectory when it finishes, assemble a batch of completed trajectories, then update the policy, with cached tool results, robust API handling, and background task curation for efficiency. The Tongyi report's version is a step-level asynchronous RL loop with separate asynchronous servers for inference and tool calls, plus tool-side result caching, timeout-and-retry, and backup APIs. Scheduling and training efficiency are deferred to L12 (RL Systems).
The Tongyi report's curves (Figure 8) show training reward rising steadily, while policy entropy rises briefly and then settles at a stable value. The slides add a caveat: these are RL training curves, and the report does not isolate each training stage.
A reward for long reports: DR Tulu's evolving rubrics
Short answers can be exact-matched; long research reports cannot. The slides list what open-ended synthesis needs to measure: coverage, relevance, factuality, and whether claims are supported by citations.
DR Tulu's answer is RLER (Reinforcement Learning with Evolving Rubrics). Each rollout produces a report, and its reward is the weighted average over rubric items:
$$ r_i = \frac{\sum_k w_k \cdot \mathrm{Judge}(c_k, y_i)}{\sum_k w_k} $$
The key is that the rubric items $c_k$ are not fixed. The slides' diagram has two layers (the IL-10 example is the slides'):
- Persistent rubrics: before training, the question is searched on the web and an LM writes rubrics from the retrieved documents; these stay for the whole run, e.g. "cites 'IL-10-engineered T cells reduced colitis severity'."
- New rubrics per instance: generated during training by contrasting good and bad rollouts for the same question. One rollout says IL-10 suppresses macrophage TNF-α via STAT3 activation; another claims an anti-inflammatory signal increases a pro-inflammatory one. The contrast yields a positive item like "states the STAT3 mechanism" and a negative item like "contains wrong claims that the anti-inflammatory cytokine upregulates…".
These rubrics go into a rubric buffer that is used both for scoring and for generating the next round of rubrics. In the paper, negative rubrics mainly target reward hacking, such as copying retrieved text verbatim to boost citation scores. The buffer has a fixed size and keeps the items whose scores vary most across rollouts for the same question. A fixed rubric stops separating good from bad once the policy has learned it; evolving rubrics update with the information the policy newly finds and the mistakes it newly makes.
The slides' ablation (the paper's Figure 6, averaged over HealthBench, ScholarQABench v2, and DeepResearch Bench; values read off the plot) starts around 50%. Random rewards reach only about 51%; initial rubrics alone reach about 59% at 2,500 steps; evolving rubrics reach about 61%. The paper's text says removing evolving rubrics costs up to 2 points, with the gap widening over training. The quality of the reward decides what RL learns — a sentence that comes back in A2.
Two more results:
- SFT as an RL warm start. From the slides' plot (the paper's Figure 5; values read off the plot): RL without SFT starts around 27%, the curve labeled Undertrained SFT starts near 46%, and the full SFT mixture gives the strongest RL results, reaching about 62% by 4,000 steps. The slide's note "even 5% SFT helps" matches the paper's text: a cold start from just 5% of the full SFT mixture already beats RL with no SFT.
- Cost. The slides' cost/performance plot puts DR Tulu-8B in the top-left corner: scores in the same band as OpenAI Deep Research and GPT-5+Search, at a per-query cost orders of magnitude lower. The paper's numbers: across four long-form benchmarks (ScholarQA-CSv2, HealthBench, ResearchQA, DeepResearch Bench) DR Tulu-8B averages 65.6, 15.6 points above Tongyi DR (the abstract writes 15.6%) and 0.7% above OpenAI DR; on ScholarQA-CSv2 it costs about USD 0.0019 per query versus about USD 1.80 for OpenAI DR, nearly three orders of magnitude less.
Retrieval: let the retriever see what the agent is thinking
Tools and dense retrieval basics
Deep research agents use two kinds of search tools: local search over your own index (lexical retrievers like BM25, embedding models) and black-box APIs (web search, browsing). Tongyi also uses Python execution, and both Tongyi and DR Tulu use scholarly search (Tongyi via Google Scholar; DR Tulu uses paper search 90% of the time on ScholarQA-CSv2).
Embedding retrieval follows the dual encoder of Dense Passage Retrieval (EMNLP 2020): encode the document collection with a document encoder and build a vector index; encode the query with a query encoder; score by inner product $s(q,d) = E_q(q)^\top E_d(d)$ and retrieve the top k. Training uses contrastive learning, pulling the query toward a relevant passage and away from negatives:
$$ L = -\log \frac{\exp s(q, d^+)}{\exp s(q, d^+) + \sum_{d^-} \exp s(q, d^-)} $$
The slides leave encoder design, negative sampling, and indexing to Advanced NLP and IR courses.
The problem: the retriever only sees the last query
The agent's context holds the original task, previous reasoning, previous queries and evidence, and the current reasoning — but only the latest query is passed to the retriever. The reasoning spells out what the agent is looking for right now and why, and the retriever never sees it.
AgentIR (arXiv:2603.04384; COLM 2026 per the slides) changes two things:
- Input: encode the current reasoning $\tau_t$ together with the query $q_t$.
- Training data (DR-Synth): standard retriever training has positives and negatives for the whole question, but each intermediate search in a trajectory wants something different, so you need local labels for "which documents help this turn." DR-Synth takes the top 50 documents from a conventional query-only retriever at this turn, puts the question's global positive documents at the front, and has an LLM run a listwise rerank using this turn's query plus the global question and answer; the top-ranked document becomes the positive and the bottom seven become hard negatives. Applied to WebShaper, this yields 5,238 training instances used to fine-tune Qwen3-Embedding-4B into AgentIR-4B.
Results
On BrowseComp-Plus, the AgentIR paper reports 68% accuracy for AgentIR-4B with Tongyi-DeepResearch, versus 52% for a conventional embedding model twice its size and 37% for BM25. The slides' scatter plot adds a second axis: AgentIR is more accurate and uses fewer search calls on average, with the same trend for gpt-oss-120B and GLM-4.7 as the agent.
The ablation separates the two changes (the paper's Table 2, accuracy, %; "+ Reasoning" prepends the reasoning to the query with no extra training):
| Agent | Query only | + Reasoning | + DR-Synth | Both (AgentIR) |
|---|---|---|---|---|
| Tongyi-DR | 48.7 | 55.5 | 59.4 | 66.3 |
| gpt-oss-120B | 47.6 | 51.3 | 59.2 | 67.0 |
| GLM-4.7 | 50.5 | 50.9 | 57.5 | 64.7 |
| Tongyi-DR (visit) | 50.2 | 54.0 | 59.5 | 68.1 |
Each change helps on its own, and combining them is best.
The last ablation asks which history helps retrieval most. All retrievers use DR-Synth; only the input changes (the paper's Table 3, accuracy, %; the paper also has a "global question" column, which the slides and this table omit):
| Agent | Query only | Prior queries | Queries + reasoning | Queries + reasoning + docs | Current reasoning (AgentIR) |
|---|---|---|---|---|---|
| Tongyi-DR | 59.4 | 63.1 | 63.1 | 60.0 | 66.3 |
| gpt-oss-120B | 59.2 | 61.9 | 64.3 | 58.7 | 67.0 |
| GLM-4.7 | 57.5 | 59.1 | 60.8 | 58.7 | 64.7 |
| Tongyi-DR (visit) | 59.5 | 63.0 | 66.3 | 61.5 | 68.1 |
The current step's reasoning is the strongest signal. Adding previously retrieved documents scores below "queries + reasoning" in all four settings, so more history is not automatically better.
The four-line summary
The final slide:
- Tasks: deep research combines many searches with evidence synthesis.
- Evaluation: assess answer quality and citation support; audit the rubrics too.
- Training: learn from demonstrations, then improve through task feedback.
- Retrieval: give retrievers the agent's information need and train for it.
Things you can try tonight
- Run a FACT-style check on your own research agent. Take one recent report, extract every "claim + citation" into its own line, open each link, and check whether the page supports the sentence. Compute citation accuracy. This tells you where your agent breaks faster than any benchmark.
- Search again with "reasoning + query." If your agent uses embedding retrieval, feed the reasoning it wrote before searching into the query encoder along with the query, and compare the top 5 results before and after. No retraining needed to see whether the signal is there.
- Audit your rubric before you use it. Write 3–5 rubric items for one research question and ask of each: does it cite something that does not exist? Does it just restate the question? Would a plausible wrong answer pass it?
Where it sits in the course
L10 is the third lecture of the Domains module, placed in the middle of the Training module: L8 covered SFT and L9 RL basics before it, while L11 covers advanced RL algorithms and L12 RL systems after it. So the modeling half of this lecture is essentially a worked application of L8 and L9 — data synthesis, rejection-filtered SFT, RLVR, and GRPO, all landed on a search agent.
It also sets up Assignment 2 (Eval). A2 asks students to write a validator for a data-visualization agent and runs into the same problems this lecture's evaluation section lists: more than one valid answer, unreliable similarity, and LM judges that need auditing. DR Tulu shows the next step: once the grader is good enough, it becomes the RL reward.
Further reading
For the series overview, see the course overview post.
Related posts on this site to read alongside the lecture:
- Deep Research Landscape: Taxonomy of 80+ Implementations, Roadmap, and Trade-offs
- How to Build a Deep Research Agent: Multi-Turn Search Planning, Conflict Resolution, and Verifiable Conclusions
- From Search Results to Reliable Citations: URL Deduplication, Source Tiers, and Claim-Source Mapping
- CS336 Lecture 16: RLVR Scales Reasoning with Verifiable Rewards, but GRPO Is Not Free PPO
- Reading Stanford CS329Z Week 7: Score Honestly, Scale Data — Midterm Checkpoint
References
Every source below was opened in full (arXiv HTML, ACL Anthology PDF, the Nature full-text page, or the original web page) and checked against the passages cited in the post.
Course materials:
- CMU 11-768 AI Agents course site
- Lecture 10 slides: Deep Research Agents (Akari Asai)
- Natural Questions data browser: the official train-set examples include "who wrote the score for the force awakens", whose document is the 2018 revision of Wikipedia's "Star Wars: The Force Awakens (soundtrack)"; its first paragraph names John Williams as the composer (checked 2026-09-29)
Assigned readings:
- OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs (arXiv:2411.14199); Nature version (the 0.79 / 0.80 rubric agreement figures are from the Nature version's Methods)
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents (arXiv:2504.12516)
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents (arXiv:2506.11763)
- Tongyi DeepResearch Technical Report (arXiv:2510.24701)
- DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research (arXiv:2511.19399)
- AgentIR: Reasoning-Aware Retrieval for Deep Research Agents (arXiv:2603.04384)
Other papers cited on the slides and used to support this post:
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (arXiv:2508.06600) (the slides cite it as "A Fair and Disentangled Evaluation Benchmark for Deep Search Agents," ACL 2026)
- FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning (arXiv:2509.13160)
- MedBrowseComp: Benchmarking Medical Deep Research and Computer Use (arXiv:2505.14963)
- ScholarSearch: Benchmarking Scholar Searching Ability of LLMs (arXiv:2506.13784)
- AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery (arXiv:2604.25256)
- A Critical Evaluation of Evaluations for Long-form Question Answering (ACL 2023)
- ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics (arXiv:2509.00496)
- DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports (arXiv:2601.08536)
- WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning (arXiv:2509.13305)
- WebDancer: Towards Autonomous Information Seeking Agency (arXiv:2505.22648)
- Scaling Agents via Continual Pre-training (arXiv:2509.13310)
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning (arXiv:2503.09516)
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (arXiv:2402.03300)
- Dense Passage Retrieval for Open-Domain Question Answering (EMNLP 2020)
Loading...