Modelplane Modelplane docs

All recipes

This document is for an older version of Modelplane.

This document applies to Modelplane v0.2 and not to the latest release v0.3.

Every validated recipe in one table: model, size, architecture, precision, and the verified hardware. Select a row for the full recipe.

ModelSizeArchPrecisionVerified onNotes
Qwen3-8B qwen8BDenseBF16EKS L4An 8.2B dense chat model on a single NVIDIA L4.
Qwen3-Coder-480B qwen480B A35BMoEBF16 / FP8EKS H200A 480B code MoE, multi-node BF16 over EFA or single-node FP8 on SGLang.
Kimi-K2 moonshotai1T A32BMoEINT4EKS H200A 1T MoE served prefill/decode disaggregated across two H200 nodes.
Llama-3.1-8B meta-llama8BDenseBF16EKSGKE L4An 8B dense chat model on a single NVIDIA L4.
GLM-4.5-Air zai-org106B A12BMoEGGUF IQ4_XSGKE A100A 106B MoE served from a GGUF checkpoint via llama.cpp on a single A100.
Nemotron-3.5-Lightning nvidia30B A3BMoENVFP4Nebius H100An open 30B MoE with 3B active parameters served NVFP4 on a single H100 on Nebius.