More like DeepSeek vs OpenAI. "How can I squeeze few more %% of performance from this old rice cooker"? vs "We only need 1T dollar to by more GPUs, and... and some marketing"
In 2 years my rice cooker will run DeepSeek v12 flashyrice.
Hehe.
More like DeepSeek vs OpenAI. "How can I squeeze few more %% of performance from this old rice cooker"? vs "We only need 1T dollar to by more GPUs, and... and some marketing"
In 2 years my rice cooker will run DeepSeek v12 flashyrice.
LLM as a solution to intelligence in saturated, Now they're more in the harness layer.
I saw just about same image of Transformer == scaling everything up in 2021. However, I still would want to credit OpenAI's researchers back then for coming up with Scaling Law. Without that nobody would have invested in that much.
If you feel like that it's because you are not paying attention, yes scaling laws matter, but there has been TONS, TONS, TONS#$@$#@ of innovation in the last 3 years. Go read the Deepseekv4 tech report, MiniMaxM3, KimiK3, KimiK2.5, etc, etc. If it was just about scale, then he with the most money wins and Google, Meta, Microsoft, Apple, Amazon would all be leading. They have more cash than all the other companies, but they can't match the creativity/innovation with cash alone.
I'm aware of this. We use Deepseek V4 flash as our local agent at our company. Its very clear we are refining gains across smaller parameter densities, however its an obvious trend even in open weight. You can improve training data, harnesses, methodologies, but boiled down most of us are sitting here especially in the local run-it-yourself crowd face palming wondering, todays model is runable, tomorrow requires twice the hardware. I'm not trying to split hairs here, but its just a meme.
Fried chicken on a salad as first healthy thing you ate all year thinking you're going down the path to revolutionize your own physical fitness starting now
Tbh you can get a similar quality of results from a ~100B model that's properly prompted and in a well made harness as you can from frontier models. Frontier models seem to be optimising for absolute dogshit input and harnesses that just work with whatever the next model is rather than being codeveloped with the model itself to get the best out of it.
This is exactly what we do at our company. We find the best performing model generally that fits on our on-site RTX 6k machines and we enhance, design, iterate, and loop to find its shortcomings and upgrade our tooling against its behavior, and we get great results from it. My point in the meme was just simply observing that everyone, even open weight are just moving to doubling parameter size which each release, with deepseek being the standout that intentionally released a powerful smaller model.
Even if we discount every other advance, the scaling alone is a wonder.
Rule 2 - Specifically opposite of local, open models