u/derspentiI've had ling-3.0-flash and glm-5.2 both in my executor slot for a few weeks. They don't split the way the benchmarks predict101 de ago. de 2026, 17:33
u/Proof_Worry9882My agent called the same tool 47 times in a row and burned my API budget - so I built a fuse for it101 de ago. de 2026, 16:56
u/fragment_meDS4 Flash full model in offload, ~600 t/s pp and 45-60 t/s tg. Terrible PP?101 de ago. de 2026, 16:44
u/laterbrehAnyone else feeling this way about the closed frontier?self.LocalLLaMA101 de ago. de 2026, 16:40
u/Proof_Worry9882i fell asleep in physics class learning about fuses and ended up building a circuit breaker for runaway llm tool loops101 de ago. de 2026, 16:36
u/esw123DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s.101 de ago. de 2026, 16:10
u/Possible_Grocery8079Is the new Laguna S 2.1 IQ4_NL GGUF stable now for local coding?101 de ago. de 2026, 15:50