43 points | by mike-the-brain 2 hours ago ago
5 comments
Amazing tech
> An agent writes in an afternoon what a chatbot writes in a month
But can you just.. not.
Your tech is so good, it speaks for itself. Don't ruin that.
Great news, has made low memory bandwidth model usage so much nicer.
I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.
vllm PR for DFlash2: https://github.com/vllm-project/vllm/pull/52816
llama.cpp PR https://github.com/ggml-org/llama.cpp/pull/27342
Amazing tech
> An agent writes in an afternoon what a chatbot writes in a month
But can you just.. not.
Your tech is so good, it speaks for itself. Don't ruin that.
Great news, has made low memory bandwidth model usage so much nicer.
I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.
vllm PR for DFlash2: https://github.com/vllm-project/vllm/pull/52816
llama.cpp PR https://github.com/ggml-org/llama.cpp/pull/27342