Great models need data you can't always move
Most machine learning starts the same way: collect a large dataset in one place, then train on it. That works fine when the data is yours to collect and move. It breaks down when the most valuable data lives on millions of individual phones, or inside hospitals that legally cannot ship patient records off-site, or across companies that don't trust each other enough to pool their raw information in a shared database.
In each of these cases, centralizing the data isn't just inconvenient, it can be illegal, impractical, or both. Healthcare data is protected by regulation in most countries. Personal data on a phone, like what someone types or how they use an app, is sensitive by nature and expensive to transmit at scale even before privacy is considered. The result is a lot of genuinely useful data that traditional centralized training simply cannot touch.
