Skip to content
All tags

#paper-reading

11 posts

CMU 10-423 Wrap-up: The Practice Exam, the HW623 Paper Presentation, and the Final Project — How the Course Checks Learning, and How to Check Yourself

Beyond its four homework assignments, CMU 10-423 checks learning four ways: 6 in-class quizzes, 2 programming tests, one comprehensive exam, and a three-person final project worth 25%. 10-623/723 students also do HW623, a paper presentation. Outside CMU you can get the practice exam with solutions (13 sections, 167 points), the HW623 handout with its 33-paper list, and the 12-page project handout. This post lays out their structure and rules and gives a self-check routine that works without peeking at the answers.

AI Engineer Interview Daily — 2026-09-26: Paper Reading

Today's Paper Reading rotation covers arXiv:2609.29875, When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression — a paper that only went up on arXiv this September. It proposes Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training-free, online method that ranks historical reasoning blocks by a frozen proxy entropy score while leaving actions, tool calls, and observations untouched. Across 260 WorkBuddyBench tasks, average reward rises from 0.699 to 0.718 while input, output, and cache-read tokens drop by 25.5%, 14.4%, and 33.3% respectively. Core concepts covered: why deleting reasoning history isn't the same as static CoT compression because it can change future actions, the nonlinear 'trajectory amplification' effect, the idea that reasoning becomes safe to forget once its conclusions have been externalized into code, files, or tool output, and how this paper's findings echo the compaction and offloading mechanisms already shipping in harnesses like Claude Code and LangChain Deep Agents. The practice question takes a reviewer's-eye view: how do you prove this method isn't just getting lucky on one benchmark's task distribution, and how would you decide when to compress context for a coding agent that runs for hours across hundreds of tool calls.

AI Engineer Interview Daily — 2026-09-19: Paper Reading

Today's Paper Reading rotation covers arXiv:2509.06917, Paper2Agent, just published in Nature in 2026. It's an automated framework that uses multiple specialist agents to analyze a research paper and its open-source code, then builds a testable, iteratively refined Model Context Protocol (MCP) server that plugs into a chat agent like Claude Code — letting you reproduce the paper's results, or answer questions the paper never asked, in natural language. The authors validate it on three case studies (AlphaGenome, ScanPy, TISSUE) and report that it surfaced a novel splicing variant linked to ADHD risk. Core concepts covered: the specialist-plus-coordinator multi-agent pattern, MCP as the standardized interface between a paper's capabilities and callable tools, how an iterative generate-test-refine loop replaces a one-shot human review, and why reproducing the paper's own results and correctly answering new queries are two separate reliability bars that both need checking. The practice question takes a reviewer's-eye view: how would you design experiments to catch an auto-generated agent that's hallucinating rather than faithfully reproducing the source method, and what guardrails would you add before shipping something like this as an internal tool.

learningguide

The Research Toolkit Trifecta: Google Scholar to Find, Moonlight to Read, CorTeX to Write

Break the research workflow into three actions — find, read, write — and pick one tool for each: Google Scholar for discovery, Moonlight AI for reading comprehension, and CorTeX for collaborative writing. All three offer free tiers and together cover the full pipeline from literature search to manuscript submission.

AI Engineer Interview Daily — 2026-09-12: Paper Reading

Today's Paper Reading rotation works through arXiv:2609.11028, BenchShield — a reward-integrity detection system for LLM-agent evaluation infrastructure. It uses static, phase-aware taint analysis to expose exploitable paths before a run, paired with runtime infrastructure-side evidence to turn "this score wasn't gamed" into a verifiable claim, lifting full-chain recall from 23-94% to 77-100% on 456 human-adjudicated trajectories, with 96% runtime detection accuracy. Core concepts cover why reward hacking is an evaluation-infrastructure problem (not just a model problem), how static and dynamic taint analysis split the work, why an "optional shortcut" that breaks no stated rule is still worth defending against, and how to read a paper's numbers by asking what evidence actually backs each claim. Paired with the companion paper BAITBENCH and the real-world Hugging Face agent-swarm security incident to make the case that reward hacking isn't just an academic concern.

AI Engineer Interview Daily — 2026-09-05: Paper Reading

Paper reading rounds don't test whether you memorized a paper's conclusion — they test whether you can break down an unfamiliar paper's problem, method, and limitations in 15-20 minutes and ask a meaningful follow-up question. Today's paper introduces invalidation contracts: attaching version stamps and cacheability hints to cached LLM-agent error-recovery suggestions, so that when server-side data drifts, the client can evict exactly the stale entries at row-level granularity instead of discarding everything or re-deriving from scratch every time. The paper's core insight is decomposing 'did this caching mechanism actually save money' into two independent variables — validity (whether the cached content is still correct, determined purely by protocol design) and compliance (whether the planner model actually adopts the suggestion on the first try, which is model-dependent: the same wire bytes get 100% first-try compliance on Claude Haiku 4.5 but can drop below 11% on Claude Sonnet 5). That decomposition itself is great interview material — it demonstrates how to split a vague performance question into two separately measurable, separately attributable factors.

AI Engineer Interview Daily — 2026-08-29: Paper Reading

A paper reading round doesn't test whether you finished the paper — it tests whether you can identify the core claim within a limited window, articulate the trade-offs behind its design choices, and raise a verifiable follow-up question. Today we use the newly published SparseRead (a token-efficient reading layer, posted to arXiv on 2026-08-23) as practice material, dissecting its regime-aware Read Gate, Reader Backends, and stateful protocol, then running a full round of 'pre-filter vs. post-hoc pruning' follow-up questions.

AI Engineer Interview Prep — 2026-08-22: Paper Reading

A paper reading round doesn't test whether you finished the paper — it tests whether you can talk about it as if you ran the research yourself: articulating the trade-offs behind key design choices, spotting gaps in the experimental design, and predicting what should come next. Today we use the newly published OneDayAgent (a long-horizon agent harness, posted to arXiv on 2026-08-04) as practice material, dissecting its task decomposition, context compression, and verify-repair mechanisms, then running through a full round of typical follow-up questions.

Paper Reading Interview Guide: How to Read, Discuss, and a Must-Read List

Paper reading interviews don't test whether you've read that specific paper — they test whether you can quickly understand a new method and identify its limitations. AI-native companies (Anthropic, OpenAI) particularly favor this format. Strategy: practice reading a paper in 30 minutes and verbally stating contribution + limitation, build your own must-read list, and practice summarizing each paper in three sentences.

aiguide

arXiv Paper Quality Assessment Guide: From Endorsement Mechanisms to a Practical Checklist

arXiv does not perform peer review, and roughly 2% of submissions are rejected. Quality judgment relies on external signals: top venue acceptance > institution + open-source reproduction > citation quality. Includes a 20-item practical checklist and a 2026 toolbox (PWC has shut down).

aideep-dive

How Do People Read arXiv Papers? A Complete Guide to Methods and Tools

Reading papers is two problems stacked together: methodology (Keshav's three-pass method, 5-10 min / 1 hour / 4-5 hours) determines how to read, and tools (arXiv HTML, alphaXiv, NotebookLM, Connected Papers, Zotero) shorten the time for each pass. AI lowers the barrier to understanding; judging correctness always stays with the human.