Skip to content
Metrale Blog
LatestEngineeringBenchmarksDesign
metrale.ai
blog.metrale.ai

Notes from the inference layer

Kernel work, measured benchmarks, and what it takes to run frontier models on hardware you own. Everything we publish is reproducible from a commit.

All EngineeringBenchmarksReleasesDesign
Design Featured

Atlas is the inference engine from desktop to hyperscaler

Speed. Security. Governance. Atlas Inference from desktop to hyperscaler. Owners of dedicated systems and renters who deliver by API both need more inference per watt.

Alexi Derkatsch Sep 11, 2026 2 min read
  • Sep 1, 2026

    DFLASH-2: the fastest single-machine numbers Atlas has produced

    66.6 tokens per second on a stock build, one DGX Spark, one stream, and every figure reproducible from a commit.

    Engineering 4 min · RS
  • Aug 31, 2026

    Seven Tenets Powering Atlas Inference Accelerated Workloads

    Atlas Inference is a free and open source LLM inference engine written from scratch in Rust. These are the seven philosophical tenets we started it on, and why we left the Python vLLM stack to do it.

    Engineering 7 min · TB
Metrale

Zero-trust inference on hardware you own. Pure Rust and CUDA, built in North Carolina.

Blog

  • Latest
  • Engineering
  • Benchmarks
  • RSS feed

Metrale

  • metrale.ai
  • Documentation
  • Benchmarks
  • Download

Community

  • GitHub
  • Discord
  • X
© 2026 Metrale · Community Edition AGPLv3
blog.metrale.ai