Intelligence may not follow division of labour
I spent 90 minutes skimming through the deepseek-coder paper. It was released in early 2024. Deepseek (& others) had started giving signals that coding is a great "agentic" use case of foundation models. It's a fascinating paper, even for a non-technical person like me, because of how intuitive, simple, and effective some of their techniques are. I was particularly impressed with them training a coding generation model in two ways - namely next-token generation (good for new code generation) and fill-in-the-middle (good for edits, debugging errors in an existing code). The technique had been applied before, but Deepseek went on to produce the kill shot of testing the right mix between these two techniques for optimal performance. Anyway, my brain went off on various tangents before I landed on one that has stuck with me for the past few days. Why didn't Deepseek - or the industry at large - use a chain of smaller, specialized models vs one 'big blob' model? This is extremely counterintuitve to me, a business generalist. Why? Because it has been drilled in my brain that specialization of labor increases output, improves quality, AND lowers cost. Our entire world is designed around the specialization of labor. And yet, here we have an example of a task where specialization not only does not work, but actively fails vs generalization. The common reasons cited are: (a) many models = Chinese whispers (lol) and errors compounding (b) faster inference (since you don't need to call multiple models) and (c) reduced time and cost to train (I can see why this mattered to Deepseek in the early days).
I will be honest that I have not intuitively grasped why this should be the case, I will continue holding this thought in my head. My inability to grasp the "big blob of intelligence" concept aside, it is clearly the direction that the foundation labs have taken over the past ~3 years. Train one big blob model and it will be the best at every task. Foundation lab direction aside, there's anecdotes to suggest this is true, my favourite one being that Bloomberg spent reasonable time and money training BloombergGPT, only to be beaten hands down by GPT4 a few weeks later at tasks BloombergGPT was designed to excel at.
Here are some 'non-duh' implications of 'big blob' > specialized:
- When creating a new piece of software (or agent) from scratch, it is stupid to NOT start with the best model, no matter the cost. The best model will be the best at anytask. When your software is released to the world, start with that best model, and keep stepping down the capability (and by definition cost) ladder until it's good enough.
- Companies or countries training or fine tuning custom models for 'their fields', 'their nuances', 'their audience' is a dumb idea. Just use the best model because it will be better than anything you can train.
- Companies gatekeeping models from accessing their data or executing their workflows are smart. Because that's the only thing the 'best models' don't have. Companies in the business of datastreams, whereby the data keeps getting continuously refreshed are even smarter. (You should assume that your stale data will enter model training datasets...at some point.).
- Does this hint that our entire theory of knowledge work whereby we have applied the specialization of labor concept from the industrial revolution is plain wrong? And that knowledge workers - and the corporations that employ them - benefit from generalization, not specialization?