LLM Benchmarking

Head to head model runs on local hardware, architecture teardowns read straight from the weights, inference tuning, and the test bench every number comes from.

LLM Benchmarking Tutorials (29)

Browse, search, and work through all available articles for this category.

29 shown · All
Context Language Models Paper Review: LLMs That Edit Their Own Context
Oct 2, 202627 min readPaper Page

Context Language Models Paper Review: LLMs That Edit Their Own Context

A plain-English review of the Context Language Models paper: how treating the context as a file lets an LLM edit its own working memory, what Suffix Cache Reuse saves, and what the results show.

Read Tutorial
ModernBERT-base Architecture Teardown: Every Layer With Real Code
Oct 2, 202630 min readArchitecture Teardowns

ModernBERT-base Architecture Teardown: Every Layer With Real Code

Open up ModernBERT-base layer by layer with real code: tokens, embeddings, LayerNorm, two-way attention, local and global layers, RoPE, GeGLU, masked-language training, pooling and a classifier head.

Read Tutorial
Qwen2.5-0.5B Architecture Teardown: Every Layer With Real Code
Oct 1, 202630 min readArchitecture Teardowns

Qwen2.5-0.5B Architecture Teardown: Every Layer With Real Code

Open up Qwen2.5-0.5B layer by layer with real code: tokens, embeddings, RMSNorm, attention from the basics to grouped-query attention, RoPE, SwiGLU, logits and the training loss.

Read Tutorial
Ming-Image-0.1-Design vs Qwen-Image 2.1: We Drew Three Kids' Physics Comics
Sep 25, 202617 min readTest Bench

Ming-Image-0.1-Design vs Qwen-Image 2.1: We Drew Three Kids' Physics Comics

We drew three 16-panel physics comics for kids with Ming-Image-0.1-Design and Qwen-Image 2.1 on one RTX 5090, and measured speed, first-try panels and lettering.

Read Tutorial
Ming-Image-0.1-Design on One RTX 5090: Designs, Comics and Photo Edits in FP8
Sep 24, 202615 min readTest Bench

Ming-Image-0.1-Design on One RTX 5090: Designs, Comics and Photo Edits in FP8

We fit inclusionAI's 46 GB design model onto one 32 GB RTX 5090 by converting it to FP8 ourselves, then made 16 designs, four comics and seven photo edits at about 8 seconds per image.

Read Tutorial
Laya 421M vs Qwen3.5 0.8B vs Qwen2.5 0.5B: Small Models for Text Classification
Sep 24, 202612 min readModel Comparisons

Laya 421M vs Qwen3.5 0.8B vs Qwen2.5 0.5B: Small Models for Text Classification

We ran the 421M Laya decision model against two LLMs of about the same size, Qwen3.5 0.8B and Qwen2.5 0.5B, on five labeling tasks on a MacBook Pro M5 Max. Laya was about twice as accurate on tickets, complaints, and intents, and faster on every task but one.

Read Tutorial
MiMo V2.6 Pro and Flash Architecture Teardown
Sep 24, 202614 min readArchitecture Teardowns

MiMo V2.6 Pro and Flash Architecture Teardown

A look inside Xiaomi's MiMo V2.6 Pro and Flash: how 128-token windows, 4-bit trained experts and two draft models fit together, and what the 9B distill actually inherits.

Read Tutorial
MiniMax image-01 vs Qwen-Image 2.1: Your Photo, a Storybook, and 48 Images
Sep 22, 202613 min readTest Bench

MiniMax image-01 vs Qwen-Image 2.1: Your Photo, a Storybook, and 48 Images

We gave MiniMax's image API and a local 8-bit Qwen-Image 2.1 the same 24 prompts at 1024x1024, with and without a reference photo, plus a five-page picture book.

Read Tutorial
Qwen-Image 2.1 on a MacBook Pro M5 Max with ComfyUI: Q8 vs Q4 Tested
Sep 21, 202612 min readTest Bench

Qwen-Image 2.1 on a MacBook Pro M5 Max with ComfyUI: Q8 vs Q4 Tested

We ran Qwen-Image 2.1 locally on a 36GB MacBook Pro M5 Max with ComfyUI at Q8 and Q4, generated blog banners, diagrams, cartoons and edits of a real photo, and measured speed, memory and swap for every image.

Read Tutorial
Qwen-Image 2.1 on an RTX 5090: Local Images for a Fraction of a Cent
Sep 21, 20268 min readTest Bench

Qwen-Image 2.1 on an RTX 5090: Local Images for a Fraction of a Cent

We ran Alibaba's new open image model at 8-bit and 4-bit on one RTX 5090, measured its speed, memory and quality, and compared its per-image cost with GPT Image 2 and Nano Banana Pro.

Read Tutorial
Bonsai 2 27B vs Qwen 3.8 27B on a MacBook Pro M5 Max
Sep 20, 202621 min readModel Comparisons

Bonsai 2 27B vs Qwen 3.8 27B on a MacBook Pro M5 Max

We ran ternary Bonsai 2 27B and its 4-bit parent Qwen 3.8 27B on the same engine on a 36GB MacBook Pro. Quality matched, Bonsai used 7 GB less memory, decoded faster, and held a 148K-token prompt that Qwen could not.

Read Tutorial
DeepSeek V4.1 Sparse Attention Explained with Pictures
Sep 7, 202629 min readArchitecture Teardowns

DeepSeek V4.1 Sparse Attention Explained with Pictures

A picture by picture walk through the sparse attention inside DeepSeek V4.1 Flash: the lightning indexer, the 512 entry pick, and how CSA2 shares one pick across many layers.

Read Tutorial
Qwen 3.8 27B Benchmark on a MacBook Pro M5 Max vs RTX 5090
Sep 4, 202618 min readInference Tuning

Qwen 3.8 27B Benchmark on a MacBook Pro M5 Max vs RTX 5090

We ran the same Qwen 3.8 27B sweep on a MacBook Pro M5 Max and an RTX 5090. The desktop decodes about four times faster, but the draft head wins on both, and only one machine stalls.

Read Tutorial
Gemini 3.8 Flash Benchmarks: Every Score Compared
Sep 3, 202613 min readModel Comparisons

Gemini 3.8 Flash Benchmarks: Every Score Compared

Every Gemini 3.8 Flash benchmark score in one table, compared with Gemini 3.7 Flash, Claude Opus 5, Claude Sonnet 5 and GPT-5.6, plus pricing and where it loses.

Read Tutorial
GLM 5.3 vs GLM 5.3 Flash Architecture Teardown
Aug 30, 202616 min readArchitecture Teardowns

GLM 5.3 vs GLM 5.3 Flash Architecture Teardown

We read the config files and every tensor shape of GLM 5.3 and GLM 5.3 Flash without downloading the weights, and worked out why the smaller model caches eight times less for each token.

Read Tutorial
I Ran 1-bit DeepSeek V4 Flash and Qwen 3.8 Flash Next on CPU (No GPU Needed)
Aug 29, 202621 min readModel Comparisons

I Ran 1-bit DeepSeek V4 Flash and Qwen 3.8 Flash Next on CPU (No GPU Needed)

We ran DeepSeek V4 Flash and Qwen 3.8 Flash Next, both at 1-bit, on 19 hard problems on one RTX 5090 desktop. The score gap is about finishing, not reasoning.

Read Tutorial
Run Qwen 3.8 Flash Next (Qwen 4) 180B on CPU Only, No GPU Needed
Aug 27, 202626 min readInference Tuning

Run Qwen 3.8 Flash Next (Qwen 4) 180B on CPU Only, No GPU Needed

We ran the 180B Qwen3.8-Flash-Next on one desktop with no GPU at all, then added a single RTX 5090, and measured speed and answer quality against Qwen 3.8 27B.

Read Tutorial
Qwen 3.8 Flash Next (Qwen 4) vs Qwen 3.8 27B Architecture Teardown
Aug 27, 202621 min readArchitecture Teardowns

Qwen 3.8 Flash Next (Qwen 4) vs Qwen 3.8 27B Architecture Teardown

We read the config files and weight maps of Qwen3.8-Flash-Next and Qwen 3.8 27B to find the four changes that let a 180B model run only 6B of its weights per token.

Read Tutorial
Granite 4.2 vs Gemma 4 vs Qwen 3.8 Benchmark
Aug 26, 202617 min readModel Comparisons

Granite 4.2 vs Gemma 4 vs Qwen 3.8 Benchmark

IBM dropped the Mamba hybrid in Granite 4.2. We measured what that costs against Gemma 4, Qwen 3.8 and Ornith 1.5 on one RTX 5090.

Read Tutorial
Ornith 1.5 9B and 35B-A3B Benchmark on RTX 5090
Aug 25, 202619 min readModel Comparisons

Ornith 1.5 9B and 35B-A3B Benchmark on RTX 5090

We benchmarked Ornith 1.5 9B and 35B-A3B against Gemma 4-26B-A4B and Qwen 3.8 27B on one RTX 5090, and found the two mixtures are nearly identical on hardware and separated entirely by training.

Read Tutorial
Qwen 3.8 27B Uncensored vs Base Abliteration Teardown
Aug 25, 202617 min readArchitecture Teardowns

Qwen 3.8 27B Uncensored vs Base Abliteration Teardown

We diffed all 866 GGUF tensors of an uncensored Qwen 3.8 27B against the original and found 131 matrices edited along one shared direction, with no measurable cost to quality.

Read Tutorial
Qwen 3.8 27B vs Qwen 3.6 27B vs Qwen 3.5 27B Teardown
Aug 19, 202619 min readArchitecture Teardowns

Qwen 3.8 27B vs Qwen 3.6 27B vs Qwen 3.5 27B Teardown

We read the GGUF headers of Qwen 3.8 27B, Qwen 3.6 27B, and Qwen 3.5 27B and ran 397 measured generations on one RTX 5090 to find where the gains between the three releases really come from.

Read Tutorial
Qwen 3.8 27B vs Muse Glimmer 30B vs Gemma 4 26B Consistency Test
Aug 19, 202625 min readModel Comparisons

Qwen 3.8 27B vs Muse Glimmer 30B vs Gemma 4 26B Consistency Test

We asked Qwen 3.8 27B, Muse Glimmer 30B, and Gemma 4 26B the same twelve questions ten times each on one RTX 5090 to measure which local model gives the same answer twice.

Read Tutorial
Qwen 3.8 27B Speed Settings on llama.cpp
Aug 16, 202622 min readInference Tuning

Qwen 3.8 27B Speed Settings on llama.cpp

We ran 45 llama.cpp configurations of Qwen 3.8 27B on one RTX 5090 to find which settings make it faster: draft depth, KV cache type, context size, and reasoning effort.

Read Tutorial
Qwen 3.8 27B vs Nemotron 3.5 vs Muse Glimmer on an RTX 5090
Aug 15, 202613 min readModel Comparisons

Qwen 3.8 27B vs Nemotron 3.5 vs Muse Glimmer on an RTX 5090

We ran Qwen 3.8 27B, Nemotron 3.5 Lightning, and Muse Glimmer on sixteen hard problems on one RTX 5090 to compare accuracy, speed, VRAM, and how each model fails.

Read Tutorial
Grok 4.6 vs Gemini 3.7 Flash on Benchmarks and Pricing
Aug 14, 202610 min readModel Comparisons

Grok 4.6 vs Gemini 3.7 Flash on Benchmarks and Pricing

Grok 4.6 vs Gemini 3.7 Flash compared on the four benchmarks both companies actually publish, plus pricing, the FrontierCode label trap, and which model fits which workload.

Read Tutorial
Nemotron 3.5 vs Muse Glimmer 30B on an RTX 5090
Aug 12, 20268 min readModel Comparisons

Nemotron 3.5 vs Muse Glimmer 30B on an RTX 5090

Hands-on benchmark of NVIDIA Nemotron 3.5 Lightning vs Meta Muse Glimmer 30B on one RTX 5090 with Ollama: generation speed, VRAM use, and long-context behavior.

Read Tutorial
Advanced LLM Benchmark Prompt Suite
May 19, 202612 min readTest Bench

Advanced LLM Benchmark Prompt Suite

A comprehensive guide on progressive benchmarking for Large Language Models using structured prompt architectures and high-quality browser-based tests.

Read Tutorial
Local LLMs 2026: Technical Reference
Mar 28, 202611 min readTest Bench

Local LLMs 2026: Technical Reference

A detailed technical guide comparing dense, sparse MoE, and hybrid SSM local LLM architectures for optimal deployment on 24GB VRAM hardware systems in 2026.

Read Tutorial