Skip to content
All tags

#vision-language-model

15 posts

CMU 10-423 HW4: Text-to-Image with a Q-Former Between a Frozen GPT-2 and a Frozen DiT — Structure, Files to Edit, and Compute

HW4 in CMU 10-423 Spring 2026 is worth 79 points. The written part covers LDMs (7), VQ-VAEs (8), CLIP (4), and VLMs through PaliGemma2 (18). The programming part (40) has you train only a Q-Former between a frozen GPT-2 and a frozen CIFAR-10 DiT, so a class-conditional diffusion model learns to take text. You write three functions, checked by 14 unit tests. The handout estimates 2–3 hours on a T4 or about 1 hour on an A100 for 25 epochs, and the captions and DiT weights come from Google Drive via download_data.sh.

CMU 10-423 L12–L13: Text-to-Image, Latent Diffusion, and Vision-Language Models

CMU 10-423 spends two lectures connecting generative models to a second modality. The second half of L12 asks how text can steer an image: three routes (GANs, autoregressive Parti, diffusion with DALL-E 2 and Imagen) lead to latent diffusion, which compresses images into an autoencoder's latent space, runs DDPM there, and reads the prompt through cross-attention. L13 goes the other way and lets a language model read images: CLIP/SigLIP or a VQ-VAE turns the image into vectors or integers for a decoder-only Transformer. What separates read-only VLMs (PaliGemma, Qwen-VL) from VLMs that can also output images (LWM, Gemini) is whether image tokens are discrete.

CS224R L17: RL for Robot Foundation Models (VLAs)

VLAs trained only with imitation learning often plateau around 80% success, while autonomous robots often need 99%+. Lecture 17 of CS224R splits "how do you improve a VLA with RL on a real robot" into three routes: recast RL as supervised learning (iterated offline RL), learn a small separate policy on the VLA's representation or diffusion noise, or learn a small policy that edits the VLA's actions. The slides call this an open research problem and describe the content as recent themes plus the speaker's opinion.

CS231N L16: Vision and Language — From CLIP's Contrastive Learning to Multimodal Foundation Models That Talk About Images

The CS231N Spring 2026 vision-and-language lecture replaces the "one model per task" approach of the first half of the course with foundation models: pre-train one model on a large, diverse dataset, then adapt it to many tasks through fine-tuning, zero-shot, or few-shot use. Three threads carry the lecture. First, CLIP: contrastive learning in both directions over 400 million image-text pairs scraped from the web, then writing class names as sentences to classify without any fine-tuning; it also has weak spots, such as failing to tell "a mug in some grass" from "some grass in a mug". Second, vision-language models from LLaVA and Flamingo to Qwen3-VL and Molmo, which feed image features into an LLM so it can look at an image and output text. Third, chaining: letting an LLM write descriptions or programs that string existing vision models together.

CME295 Lecture 9: Transformers Leave Text Behind, and LLMs Stop Writing Left to Right

The last CME295 lecture packs 128 slides into three parts: an eight-picture recap of the quarter, how Transformers handle images (ViT and two ways to build a VLM), and masked diffusion LLMs that emit several tokens per step, followed by what comes next in research and applications. It is not on the exam; the 2026 edition turns diffusion LLMs into a lecture of their own and refocuses Lecture 9 on multimodality.

Commercial Document Parsing APIs Compared: Specialized Parsers, General VLMs, and the Big Three Clouds

Three routes to commercial document parsing: specialized parsers (Cohere Parse at $1.50/k pages, LlamaParse Agentic Plus at 90.2% on ParseBench), Big Three cloud prebuilts (Azure/Google/AWS for structured field extraction), and general-purpose VLMs (Fable 5.1 scores 78.92 on ParseBench and crushes specialized parsers on charts, but costs 3–16× more and hallucinates). At 100K pages/month, plain OCR runs ~$150 across providers; add tables and AWS jumps to $1,500, Claude Sonnet 5 to $900. The first question isn't 'which is most accurate' — it's 'do you need transcription or comprehension?'

Agentic Parsing: Letting Agents Decide How to Parse Documents

Traditional document parsing runs a fixed pipeline regardless of input, but contracts, financial reports, and technical manuals each need different strategies. Agentic Parsing lets LLM agents observe a document and dynamically choose tools — AgenticOCR parses only the regions that matter (70%+ visual token savings), and ParseBench shows even the best method scores only 84.9% across 2,000 enterprise pages. No silver bullet.

CS224N Lecture 17: An Official Reading Map for Multimodality

Lecture 17 is Luke Zettlemoyer's multimodality guest session, but the site publishes no slides or agenda. Its official readings establish three routes: visual reasoning workspaces, early-fusion token models, and text autoregression with image diffusion.

Stanford CS224V Lecture 12: CHURRO Makes Multilingual Historical Documents Searchable

CHURRO represents full-page text, layout, and metadata in HDML, unifies multilingual historical data for a page-level VLM, and connects extraction to HistoryGenie for searchable, conversational archives.

CS336 Lecture 17: Multimodal Models Turn Images into Tokens, Then Reconcile Semantics with Detail

Lecture 17 organizes CLIP/SigLIP, LLaVA, Qwen-VL, and Chameleon into three paths: contrastive encoders learn semantics, vision-encoder/projector/LM stacks provide understanding, and discrete image tokens enable generation. Resolution, token budgets, and modality balance constrain them all.

aideep-dive

Multimodal Models, First Half of 2026: Native Fusion vs. Bolted-On Vision, and Why Leaderboards Contradict Each Other

Pure image understanding has flattened out — four frontier models all clear 80% on MMMU-Pro within 3 points of each other. The real differentiation is video, long-document OCR, and realtime speech, each with a different leader. But the most useful lesson from assembling these rankings is that two credible sources named different Video-MME leaders more than 10 points apart — and that July and August each turned the field over again.

Scanned PDF Benchmark: How Did 10 Parsers Handle Graduate Entrance Exams?

I tested 10 open-source PDF parsing tools on four scanned NTU graduate entrance exams. VLM-based tools—Firecrawl, MinerU 3.4, and Marker v2—overwhelmingly beat conventional OCR on formulas and code, but installation was the real barrier: MinerU's old package name creates dependency hell, Marker's first model download takes 10 minutes, and PaddleOCR needs a separate engine. In practice, use RapidOCR for screening and MinerU or Firecrawl for close inspection.

The Parsing Layer: When Structure Must Be Inferred — and Licensing Is the Real Selection Axis

Scans and complex layouts leave you no choice but to infer structure with a model. But the technical gap between MinerU, Marker, and Docling is far smaller than the licensing gap — MinerU needs a separate license past $20M monthly revenue, Marker's model weights need payment past a funding threshold, and only Docling is cleanly MIT. Read the LICENSE before the benchmark.

Midscene.js: Betting on Pure Vision for Cross-Platform UI Automation

An MIT-licensed open-source UI automation framework from ByteDance. UI actions rely solely on feeding screenshots to a vision-language model, with no DOM parsing. A single JS API works across Web / Android / iOS / desktop. The trade-offs: each step is slower and more token-expensive, and everything hinges on the model's grounding ability. Note that Midscene retired MCP after 1.9.8 in favour of Skills + CLI.

aideep-dive

DeepSeek-OCR: The 10x Compression Experiment That Turns Long Context into Images

DeepSeek-OCR's paper is titled Contexts Optical Compression -- OCR is just the means; what it actually validates is that 'rendering text as images and feeding them to a VLM' achieves 10x compression at 97% accuracy. This is a qualitative shift for long-context LLM and RAG token costs.