Wafer Blog Posts

Emilio AndereSeptember 1, 2026

Wafer Raises $40M Series A to Build AI That Optimizes AI

The round was co-led by Marathon and Chemistry to accelerate Wafer's vision of AI that continuously optimizes AI inference.

Ian YeJuly 31, 2026

Is memory the moat?

Running Kimi K3 at ~952 tok/s/node, AMD continues to prove its case as the winner in performance per dollar.

Ian YeJuly 3, 2026

Performance per dollar is getting faster and cheaper

How we served GLM5.2 on AMD MI355X at 2626 tok/s/node and 213 tok/s single stream at over 2x lower cost than Blackwell.

Balaji Varadarajan and Wafer TeamJune 10, 2026

The Inference Alpha: Maximizing Frontier Models on AMD

How DigitalOcean and Wafer unlock order-of-magnitude inference speedups on AMD GPUs for Kimi 2.5, DeepSeek V3.2, and GLM-5 through deep kernel and systems engineering.

Ian YeMay 19, 2026

Achieving Heterogeneous Compute One Kernel at a Time

How custom kernels pushed our AMD MI355X deployment from a tuned baseline to leading Qwen3.5 397B throughput.

John HahnMay 13, 2026

Quantizing Kimi K2.6 to NVFP4 for Blackwell Inference

Wafer and Parasail released wafer-ai/Kimi-K2.6-NVFP4: Blackwell NVFP4 weights for production Kimi K2.6 inference.

Danny Willow LiuMay 11, 2026

Qwen3.6-35B-A3B on AMD MI355X: Being The Fastest With ATOM on AMD

ATOM on 8×MI355X leads public Qwen3.6-35B-A3B on Artificial Analysis decode and sustains ~15k tok/s per node at production latency.

Emilio AndereApril 14, 2026

Announcing our Seed Round

Wafer has raised $4 million in seed funding led by Fifty Years to build AI that optimizes AI infrastructure.

Steven ArellanoMarch 17, 2026

Where Did My Microseconds Go?

In our NVFP4 KernelArena suite, cpp_extension.load() hid a blind spot: JIT and CPU migration inflated kernel launch times.

Emilio AndereMarch 12, 2026

A Field Guide to Reward Hacking in AI Kernel Generation

Ten ways LLMs game GPU kernel benchmarks: timer tricks, garbage reads, caching, and the checks that catch them.

Steven ArellanoMarch 11, 2026

Introducing KernelArena

Benchmark AI-generated GPU kernels, with first results on WaferBench NVFP4 (B200) and KernelBench HIP (MI300X).

Wafer TeamFebruary 10, 2026

Trace Compare: Compare vLLM traces across platforms

1:1 kernel mappings across providers. Diff huge vLLM traces fast with clean prefill vs decode.

Wafer TeamFebruary 5, 2026

Workspaces: GPU Compute for Your Coding Agent

Give your AI coding assistant direct GPUs, with no manual SSH, Docker, or infra babysitting.

Wafer TeamFebruary 3, 2026

Cloud Compiler Analyzer (PTX/SASS) Inside Your IDE

Cloud CUDA builds with PTX/SASS, PyTorch headers, and VS Code integration. No local CUDA install.

Wafer TeamFebruary 1, 2026

Nordlys Labs: 8x Faster Routing with Wafer-Guided Kernel Optimization

A non-kernel expert hit 8× on latency-critical CUDA clustering with Wafer profile-guided optimization.

Wafer TeamJanuary 30, 2026

Profile-Guided GPU Kernel Optimization

Profiling in our CLI broke a theory-only plateau: 11.65× on the Kimi Delta Attention kernel.

Wafer TeamJanuary 29, 2026

The Year of the LLM GPU Kernel Engineer

Our agent optimized AMD's topk_sigmoid kernel: 9× over PyTorch, step by step.

Wafer TeamJanuary 27, 2026

Case Study: A 104x (?) Speedup on KernelBench

A fused kernel claimed 104× speedup while reading garbage, and passed checks until we added a determinism guard.

Wafer TeamJanuary 23, 2026

Which models are the most HIP?

Frontier models wrote HIP kernels for KernelBench on MI300X. We measured which ones were correct and how fast.

Wafer TeamJanuary 16, 2026

GPU Docs: Now Available on the Web

The GPU docs tool from our IDE extension, now a standalone web app.

Wafer TeamJanuary 13, 2026

Introducing ROCprofiler Compute: AMD GPU Profiling in Your IDE

AMD profiling in VS Code and Cursor: metrics, roofline, and kernel stats without leaving the editor.

Wafer TeamJanuary 8, 2026

Introducing Wafer's Built-in Perfetto Trace Viewer

Open Chrome trace JSON in your IDE with Perfetto: timeline, flamegraphs, SQL, metrics.

Wafer TeamDecember 19, 2025

Introducing the Wafer Extension for VS Code and Cursor

Profiling (NCU), compiler explorer, and GPU docs: your GPU stack inside the editor.

Wafer TeamJuly 14, 2025

Introducing Chip Benchmark: Hardware-Centric Performance Insights for AI Workloads

Chip Benchmark is our open suite for open-weight LLMs across accelerators, so you can pick hardware with real numbers.

Wafer TeamJuly 14, 2025

Unlocking AMD MI300X for High-Throughput, Low-Cost LLM Inference

MI300X packs 192GB HBM3 and 5.3 TB/s bandwidth, often skipped for LLM inference. We show quantization and tuning.