u/derspentiI've had ling-3.0-flash and glm-5.2 both in my executor slot for a few weeks. They don't split the way the benchmarks predict10Aug 1, 2026, 5:33 PM
u/Proof_Worry9882My agent called the same tool 47 times in a row and burned my API budget - so I built a fuse for it10Aug 1, 2026, 4:56 PM
u/fragment_meDS4 Flash full model in offload, ~600 t/s pp and 45-60 t/s tg. Terrible PP?10Aug 1, 2026, 4:44 PM
u/laterbrehAnyone else feeling this way about the closed frontier?self.LocalLLaMA10Aug 1, 2026, 4:40 PM
u/Proof_Worry9882i fell asleep in physics class learning about fuses and ended up building a circuit breaker for runaway llm tool loops10Aug 1, 2026, 4:36 PM
u/esw123DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s.10Aug 1, 2026, 4:10 PM
u/Possible_Grocery8079Is the new Laguna S 2.1 IQ4_NL GGUF stable now for local coding?10Aug 1, 2026, 3:50 PM