Wafer Raises $40M Series A to Build AI That Optimizes AI
The round was co-led by Marathon and Chemistry to accelerate Wafer's vision of AI that continuously optimizes AI inference.

Wafer Raises $40M Series A to Build AI That Optimizes AI
The round was co-led by Marathon and Chemistry to accelerate Wafer's vision of AI that continuously optimizes AI inference.

Is memory the moat?
Running Kimi K3 at ~952 tok/s/node, AMD continues to prove its case as the winner in performance per dollar.

Performance per dollar is getting faster and cheaper
How we served GLM5.2 on AMD MI355X at 2626 tok/s/node and 213 tok/s single stream at over 2x lower cost than Blackwell.

The Inference Alpha: Maximizing Frontier Models on AMD
How DigitalOcean and Wafer unlock order-of-magnitude inference speedups on AMD GPUs for Kimi 2.5, DeepSeek V3.2, and GLM-5 through deep kernel and systems engineering.

Achieving Heterogeneous Compute One Kernel at a Time
How custom kernels pushed our AMD MI355X deployment from a tuned baseline to leading Qwen3.5 397B throughput.

Quantizing Kimi K2.6 to NVFP4 for Blackwell Inference
Wafer and Parasail released wafer-ai/Kimi-K2.6-NVFP4: Blackwell NVFP4 weights for production Kimi K2.6 inference.

Qwen3.6-35B-A3B on AMD MI355X: Being The Fastest With ATOM on AMD
ATOM on 8×MI355X leads public Qwen3.6-35B-A3B on Artificial Analysis decode and sustains ~15k tok/s per node at production latency.

Announcing our Seed Round
Wafer has raised $4 million in seed funding led by Fifty Years to build AI that optimizes AI infrastructure.

Where Did My Microseconds Go?
In our NVFP4 KernelArena suite, cpp_extension.load() hid a blind spot: JIT and CPU migration inflated kernel launch times.

A Field Guide to Reward Hacking in AI Kernel Generation
Ten ways LLMs game GPU kernel benchmarks: timer tricks, garbage reads, caching, and the checks that catch them.

Introducing KernelArena
Benchmark AI-generated GPU kernels, with first results on WaferBench NVFP4 (B200) and KernelBench HIP (MI300X).

Trace Compare: Compare vLLM traces across platforms
1:1 kernel mappings across providers. Diff huge vLLM traces fast with clean prefill vs decode.

Workspaces: GPU Compute for Your Coding Agent
Give your AI coding assistant direct GPUs, with no manual SSH, Docker, or infra babysitting.

Cloud Compiler Analyzer (PTX/SASS) Inside Your IDE
Cloud CUDA builds with PTX/SASS, PyTorch headers, and VS Code integration. No local CUDA install.

Nordlys Labs: 8x Faster Routing with Wafer-Guided Kernel Optimization
A non-kernel expert hit 8× on latency-critical CUDA clustering with Wafer profile-guided optimization.

Profile-Guided GPU Kernel Optimization
Profiling in our CLI broke a theory-only plateau: 11.65× on the Kimi Delta Attention kernel.

The Year of the LLM GPU Kernel Engineer
Our agent optimized AMD's topk_sigmoid kernel: 9× over PyTorch, step by step.

Case Study: A 104x (?) Speedup on KernelBench
A fused kernel claimed 104× speedup while reading garbage, and passed checks until we added a determinism guard.

Which models are the most HIP?
Frontier models wrote HIP kernels for KernelBench on MI300X. We measured which ones were correct and how fast.

GPU Docs: Now Available on the Web
The GPU docs tool from our IDE extension, now a standalone web app.

Introducing ROCprofiler Compute: AMD GPU Profiling in Your IDE
AMD profiling in VS Code and Cursor: metrics, roofline, and kernel stats without leaving the editor.

Introducing Wafer's Built-in Perfetto Trace Viewer
Open Chrome trace JSON in your IDE with Perfetto: timeline, flamegraphs, SQL, metrics.

Introducing the Wafer Extension for VS Code and Cursor
Profiling (NCU), compiler explorer, and GPU docs: your GPU stack inside the editor.

Introducing Chip Benchmark: Hardware-Centric Performance Insights for AI Workloads
Chip Benchmark is our open suite for open-weight LLMs across accelerators, so you can pick hardware with real numbers.

Unlocking AMD MI300X for High-Throughput, Low-Cost LLM Inference
MI300X packs 192GB HBM3 and 5.3 TB/s bandwidth, often skipped for LLM inference. We show quantization and tuning.