
LLM Benchmarking
Head to head model runs on local hardware, architecture teardowns read straight from the weights, inference tuning, and the test bench every number comes from.
LLM Benchmarking Tutorials (29)
Browse, search, and work through all available articles for this category.
Context Language Models Paper Review: LLMs That Edit Their Own Context
A plain-English review of the Context Language Models paper: how treating the context as a file lets an LLM edit its own working memory, what Suffix Cache Reuse saves, and what the results show.
Read TutorialModernBERT-base Architecture Teardown: Every Layer With Real Code
Open up ModernBERT-base layer by layer with real code: tokens, embeddings, LayerNorm, two-way attention, local and global layers, RoPE, GeGLU, masked-language training, pooling and a classifier head.
Read TutorialQwen2.5-0.5B Architecture Teardown: Every Layer With Real Code
Open up Qwen2.5-0.5B layer by layer with real code: tokens, embeddings, RMSNorm, attention from the basics to grouped-query attention, RoPE, SwiGLU, logits and the training loss.
Read TutorialMing-Image-0.1-Design vs Qwen-Image 2.1: We Drew Three Kids' Physics Comics
We drew three 16-panel physics comics for kids with Ming-Image-0.1-Design and Qwen-Image 2.1 on one RTX 5090, and measured speed, first-try panels and lettering.
Read TutorialMing-Image-0.1-Design on One RTX 5090: Designs, Comics and Photo Edits in FP8
We fit inclusionAI's 46 GB design model onto one 32 GB RTX 5090 by converting it to FP8 ourselves, then made 16 designs, four comics and seven photo edits at about 8 seconds per image.
Read TutorialLaya 421M vs Qwen3.5 0.8B vs Qwen2.5 0.5B: Small Models for Text Classification
We ran the 421M Laya decision model against two LLMs of about the same size, Qwen3.5 0.8B and Qwen2.5 0.5B, on five labeling tasks on a MacBook Pro M5 Max. Laya was about twice as accurate on tickets, complaints, and intents, and faster on every task but one.
Read TutorialMiMo V2.6 Pro and Flash Architecture Teardown
A look inside Xiaomi's MiMo V2.6 Pro and Flash: how 128-token windows, 4-bit trained experts and two draft models fit together, and what the 9B distill actually inherits.
Read TutorialMiniMax image-01 vs Qwen-Image 2.1: Your Photo, a Storybook, and 48 Images
We gave MiniMax's image API and a local 8-bit Qwen-Image 2.1 the same 24 prompts at 1024x1024, with and without a reference photo, plus a five-page picture book.
Read TutorialQwen-Image 2.1 on a MacBook Pro M5 Max with ComfyUI: Q8 vs Q4 Tested
We ran Qwen-Image 2.1 locally on a 36GB MacBook Pro M5 Max with ComfyUI at Q8 and Q4, generated blog banners, diagrams, cartoons and edits of a real photo, and measured speed, memory and swap for every image.
Read TutorialQwen-Image 2.1 on an RTX 5090: Local Images for a Fraction of a Cent
We ran Alibaba's new open image model at 8-bit and 4-bit on one RTX 5090, measured its speed, memory and quality, and compared its per-image cost with GPT Image 2 and Nano Banana Pro.
Read TutorialBonsai 2 27B vs Qwen 3.8 27B on a MacBook Pro M5 Max
We ran ternary Bonsai 2 27B and its 4-bit parent Qwen 3.8 27B on the same engine on a 36GB MacBook Pro. Quality matched, Bonsai used 7 GB less memory, decoded faster, and held a 148K-token prompt that Qwen could not.
Read TutorialDeepSeek V4.1 Sparse Attention Explained with Pictures
A picture by picture walk through the sparse attention inside DeepSeek V4.1 Flash: the lightning indexer, the 512 entry pick, and how CSA2 shares one pick across many layers.
Read TutorialQwen 3.8 27B Benchmark on a MacBook Pro M5 Max vs RTX 5090
We ran the same Qwen 3.8 27B sweep on a MacBook Pro M5 Max and an RTX 5090. The desktop decodes about four times faster, but the draft head wins on both, and only one machine stalls.
Read TutorialGemini 3.8 Flash Benchmarks: Every Score Compared
Every Gemini 3.8 Flash benchmark score in one table, compared with Gemini 3.7 Flash, Claude Opus 5, Claude Sonnet 5 and GPT-5.6, plus pricing and where it loses.
Read TutorialGLM 5.3 vs GLM 5.3 Flash Architecture Teardown
We read the config files and every tensor shape of GLM 5.3 and GLM 5.3 Flash without downloading the weights, and worked out why the smaller model caches eight times less for each token.
Read TutorialI Ran 1-bit DeepSeek V4 Flash and Qwen 3.8 Flash Next on CPU (No GPU Needed)
We ran DeepSeek V4 Flash and Qwen 3.8 Flash Next, both at 1-bit, on 19 hard problems on one RTX 5090 desktop. The score gap is about finishing, not reasoning.
Read TutorialRun Qwen 3.8 Flash Next (Qwen 4) 180B on CPU Only, No GPU Needed
We ran the 180B Qwen3.8-Flash-Next on one desktop with no GPU at all, then added a single RTX 5090, and measured speed and answer quality against Qwen 3.8 27B.
Read TutorialQwen 3.8 Flash Next (Qwen 4) vs Qwen 3.8 27B Architecture Teardown
We read the config files and weight maps of Qwen3.8-Flash-Next and Qwen 3.8 27B to find the four changes that let a 180B model run only 6B of its weights per token.
Read TutorialGranite 4.2 vs Gemma 4 vs Qwen 3.8 Benchmark
IBM dropped the Mamba hybrid in Granite 4.2. We measured what that costs against Gemma 4, Qwen 3.8 and Ornith 1.5 on one RTX 5090.
Read TutorialOrnith 1.5 9B and 35B-A3B Benchmark on RTX 5090
We benchmarked Ornith 1.5 9B and 35B-A3B against Gemma 4-26B-A4B and Qwen 3.8 27B on one RTX 5090, and found the two mixtures are nearly identical on hardware and separated entirely by training.
Read TutorialQwen 3.8 27B Uncensored vs Base Abliteration Teardown
We diffed all 866 GGUF tensors of an uncensored Qwen 3.8 27B against the original and found 131 matrices edited along one shared direction, with no measurable cost to quality.
Read TutorialQwen 3.8 27B vs Qwen 3.6 27B vs Qwen 3.5 27B Teardown
We read the GGUF headers of Qwen 3.8 27B, Qwen 3.6 27B, and Qwen 3.5 27B and ran 397 measured generations on one RTX 5090 to find where the gains between the three releases really come from.
Read TutorialQwen 3.8 27B vs Muse Glimmer 30B vs Gemma 4 26B Consistency Test
We asked Qwen 3.8 27B, Muse Glimmer 30B, and Gemma 4 26B the same twelve questions ten times each on one RTX 5090 to measure which local model gives the same answer twice.
Read TutorialQwen 3.8 27B Speed Settings on llama.cpp
We ran 45 llama.cpp configurations of Qwen 3.8 27B on one RTX 5090 to find which settings make it faster: draft depth, KV cache type, context size, and reasoning effort.
Read TutorialQwen 3.8 27B vs Nemotron 3.5 vs Muse Glimmer on an RTX 5090
We ran Qwen 3.8 27B, Nemotron 3.5 Lightning, and Muse Glimmer on sixteen hard problems on one RTX 5090 to compare accuracy, speed, VRAM, and how each model fails.
Read TutorialGrok 4.6 vs Gemini 3.7 Flash on Benchmarks and Pricing
Grok 4.6 vs Gemini 3.7 Flash compared on the four benchmarks both companies actually publish, plus pricing, the FrontierCode label trap, and which model fits which workload.
Read TutorialNemotron 3.5 vs Muse Glimmer 30B on an RTX 5090
Hands-on benchmark of NVIDIA Nemotron 3.5 Lightning vs Meta Muse Glimmer 30B on one RTX 5090 with Ollama: generation speed, VRAM use, and long-context behavior.
Read TutorialAdvanced LLM Benchmark Prompt Suite
A comprehensive guide on progressive benchmarking for Large Language Models using structured prompt architectures and high-quality browser-based tests.
Read TutorialLocal LLMs 2026: Technical Reference
A detailed technical guide comparing dense, sparse MoE, and hybrid SSM local LLM architectures for optimal deployment on 24GB VRAM hardware systems in 2026.
Read Tutorial