$ ls ~/recipes/
Recipes
How I serve models on DGX Spark. Every config here is real and in use: the lmswitch recipes I switch between day to day, and the reproducible recipe behind every run on the DGX Spark bench.
$ ls ~/ai-models/*.yaml
lmswitch configs from jvr0x/ai-models (llama.cpp, vLLM, dual-Spark vLLM and ExLlamaV3 jobs). Each links to its YAML on GitHub.
DeepSeek-V4-Flash 0731 Abliterated 1M [DUAL 2x SPARK]
Keys abliterated DeepSeek-V4-Flash-0731 - TP=2 across spark + gigabyte over CX7 (1M ctx).
DeepSeek-V4-Flash 0731 1M [DUAL 2x SPARK]
DeepSeek V4 Flash 0731 - TP=2 across spark + gigabyte over CX7 (1M ctx).
DeepSeek-V4-Flash DSpark 1M [DUAL 2x SPARK]
DeepSeek V4 Flash DSpark - TP=2 across spark + gigabyte over CX7 (1M ctx).
Gemma-4-12B-IT NVFP4 (Unsloth)
Gemma-4-12B-IT NVFP4 (Unsloth) - small-footprint vLLM profile.
Gemma-4-12B-IT QAT Q4 MTP
Gemma-4-12B-IT QAT GGUF (Unsloth) - vision + MTP speculative decoding.
Mistral-Medium-3.5-128B NVFP4 [BENCH]
lmswitch BENCH PROFILE of Mistral-Medium-3.5-128B NVFP4 - dense 128B via vLLM.
Nemotron-3-Super-120B NVFP4 [BENCH]
lmswitch BENCH PROFILE of Nemotron-3-Super-120B-A12B NVFP4 via vLLM (stock image).
Nemotron-3-Super-120B NVFP4
lmswitch recipe for NVIDIA Nemotron-3-Super-120B-A12B-NVFP4 on 1× DGX Spark.
Nemotron-Labs-3-Puzzle-75B-A9B NVFP4
lmswitch recipe for NVIDIA Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4 on 1× DGX Spark.
Ornith-1.0-35B NVFP4 (AEON) [BENCH]
lmswitch BENCH PROFILE of Ornith-1.0-35B-AEON NVFP4 - hybrid Mamba/Attention.
Qwen3.6-27B NVFP4 (MiaAI-Lab repo config)
Faithful port of MiaAI-Lab/Qwen3.6-27B-NVFP4-vLLM start.sh, as a lmswitch recipe.
Qwen3.6-27B NVFP4 (Unsloth) [BENCH]
lmswitch BENCH PROFILE of qwen3.6-27b-nvfp4-unsloth - local NVFP4 dense hybrid VLM.
Qwen3.6-27B NVFP4 [UNSLOTH]
lmswitch recipe for Qwen3.6-27B-NVFP4 (unsloth) on 1× DGX Spark.
Qwen3.6-35B-A3B NVFP4-Fast [UNSLOTH]
lmswitch recipe for Qwen3.6-35B-A3B-NVFP4-Fast (unsloth) on 1× DGX Spark.
Qwen3.6-35B-A3B NVFP4 Fast (Unsloth) [BENCH]
lmswitch BENCH PROFILE of qwen3.6-35b-nvfp4-unsloth-fast - local NVFP4 MoE hybrid.
RED-SNOW-5.3-FLASH EXL3-SAGE 4.91bpw [DUAL 2x SPARK job]
lmswitch job for RED-SNOW-5.3-FLASH EXL3 SAGE 4.91bpw across spark + gigabyte.
$ ls ~/forks/
DeepSeek-V4.1-Flash EXL3 on one DGX Spark
Serving DeepSeek-V4.1-Flash as EXL3 on a single DGX Spark, about 17.5 tok/s with the DSpark drafter.
$ ls ~/dgx-spark-bench/recipes/
One entry per measured config on the bench dashboard, linked to the recipe that produced it. Where a bench config is the same file as an lmswitch recipe, it links to that too.
DeepSeek-V4-Flash (0731)
Upstream recipe by MiaAI-Lab, measured on the bench.
GLM-5.3-Flash
Upstream recipe by vcruz305, measured on the bench.
Kolibri-1
Upstream recipe by Aleph-Alpha, measured on the bench.
Qwen3.8-Flash-Next
Upstream recipe by MiaAI-Lab, measured on the bench.