Skip to content
All tags

#perplexity

7 posts

NTHU Hung-Yu Kao NLP, Week 2: Word Embeddings and Language Models, from Counting N-grams to RNNs

Week 2 of Hung-Yu Kao's NLP course at NTHU is a 62-page deck about predicting the next word. The first part covers statistical language models with a bigram count table, add-one smoothing, and perplexity, then sparse vectors through a PPMI example with cherry and digital. The middle returns to Word2Vec's negative sampling and uses the Chinese word for 'apple' to show why contextualized embeddings are needed. The last part derives RNNs from three weaknesses of feedforward networks and shows how they handle NER, sentence classification, and stacked and bidirectional variants.

Pricing Watch | Perplexity Retires Sonar API Today, Moves to Token-Plus-Tool-Call Billing on the Agent API

Perplexity's pricing docs and migration guide confirm that Sonar Chat Completions (sonar, sonar-pro, sonar-reasoning-pro, sonar-deep-research) stops being supported on 2026-09-27, fully replaced by the Agent API. The old model billed model token price plus a flat request fee tiered by search depth ($5-$12 per 1,000 requests); the new model bills whatever third-party model token price you pick (e.g. gpt-5.6-luna at $0.20/1M input) plus per-tool-call fees (web_search at $0.0025, fetch_url at $0.0005). Using Perplexity's own representative usage figures, Sonar-to-fast comes out about 49% cheaper, Sonar Pro-to-low about 84% cheaper, and Sonar Reasoning Pro-to-medium about 71% cheaper. But sonar-pro and sonar-reasoning-pro stop being routable outright on most third-party gateways — only the base sonar model gets auto-migrated to the Agent API, and every other tier requires a manual switch to the new preset system.

Commercial Landscape: OpenAI, Perplexity, Gemini, Claude, Grok

By 2026, the deep research commercial market has differentiated: OpenAI is comprehensive, Perplexity is fast, Gemini integrates ecosystems, Claude reasons deeply, Grok is real-time. This article compares each product's differences—not who is best, but who fits your scenario.

How to Read a Model's Report Card: Benchmarks, Arena Elo, and the Traps Behind the Numbers

Benchmark scores in model releases have three common traps: cherry-picking (only showing wins), contamination (test data leaking into training), and saturation (when everyone scores 90%+, the benchmark stops being useful). The most manipulation-resistant signal is Chatbot Arena's Elo ranking — real humans, blind voting, uncontrolled questions.

How a Model Knows It's Wrong: Loss Functions and Cross-Entropy

Every time a model predicts the next token, it assigns a probability to every candidate word. A loss function measures how far that probability distribution is from the correct answer — the further off, the higher the loss, the more the model knows it got it wrong. Cross-entropy is the standard formula; perplexity is its human-readable translation.

techdeep-dive

Bumblebee: A Design Teardown of Perplexity's Read-Only Supply Chain Endpoint Scanner

A Go read-only scanner open-sourced by Perplexity in May 2026 (v0.1.1, zero non-stdlib dependencies). It inventories npm/PyPI/Go/RubyGems/Composer/MCP/editor and browser extensions into NDJSON, matches against a custom exposure catalog, and answers the question 'which machines in my fleet are currently affected' the moment a supply chain incident hits. It deliberately never invokes any package manager and is not an EDR.

Is Your JSON-LD Invisible to AI Search Engines? A Pipeline Breakdown and AEO/GEO Strategy

Different AI engines process web pages in vastly different ways. Some only read the body; others rely on pre-built indexes. JSON-LD and schema markup are not universally effective — body content quality and structure are the only cross-platform foundations that hold.