A Tesla AI senior principal engineer named Cai Yunda has a message for anyone who thinks machine learning is all about training models: you're wrong. According to Cai, the popular assumption that ML projects spend 99% of their time on training is backwards. In reality, only about 2% of effort goes into model parameter training. The rest? 50% is evaluation and testing, 40% is cleaning data, and 8% is system integration.
https://twitter.com/yunta_tsai/status/2068364559698780520
Cai emphasizes that data cleaning and evaluation fundamentally define the ceiling of what AI can learn. If the raw data has fuzzy definitions or inconsistent labels, you're injecting noise from the start. No amount of algorithmic wizardry or hyperparameter tuning can erase that background noise — the model can't correct its own flawed textbook. The final accuracy is bounded by the effective information content in the data, a hard limit rooted in Shannon's coding theory.
To keep data standards consistent, Cai says he spends his days reexamining the definitions and classification systems behind data concepts — what he calls ontology — and even revisiting historical labels. Whether you're setting rules for reinforcement learning or fine-tuning with precise annotations, the deciding factor is always data quality and evaluation rigor, not the model architecture itself.