Rendered at 01:33:37 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
ilc 2 hours ago [-]
Watch the video carefully. DFlash2's tool call fails on python syntax.
Usually models in this class nail things like that 1 shot, which the other side did.
I don't know the cause. It may be nothing. But I'd like to see the model doing something where its path is a bit more constrained, to help out rule out such oddities.
zackangelo 1 hours ago [-]
DFlash is lossless so this would be a bug in the implementation if it is indeed a regression against the target model.
liuliu 48 minutes ago [-]
Only if you do greedy sampling. With probabilisitic sampling (categorical sampling), you will end up with different trajectory just “mathematically equivalent”.
hypfer 4 hours ago [-]
Amazing tech
> An agent writes in an afternoon what a chatbot writes in a month
But can you just.. not.
Your tech is so good, it speaks for itself. Don't ruin that.
adefa 4 hours ago [-]
I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.
sarjann 4 hours ago [-]
Great news, has made low memory bandwidth model usage so much nicer.
Usually models in this class nail things like that 1 shot, which the other side did.
I don't know the cause. It may be nothing. But I'd like to see the model doing something where its path is a bit more constrained, to help out rule out such oddities.
> An agent writes in an afternoon what a chatbot writes in a month
But can you just.. not.
Your tech is so good, it speaks for itself. Don't ruin that.