Tesla AI senior staff engineer Cai Yunda says there's a common misconception that machine learning projects spend 99% of their time on training. In reality, only about 2% of the work goes into model parameter training. The real split? 50% on evaluation and testing, 40% on data cleaning, and the remaining 8% on system integration.
https://twitter.com/yunta_tsai/status/2068364559698780520
Cai emphasizes that data cleaning and evaluation determine the absolute ceiling of what an AI can learn. If the raw data has fuzzy definitions or inconsistent labels, it introduces noise from the very start. No amount of algorithmic magic or hyperparameter tuning can eliminate background noise — the model can't correct its own faulty textbook. The final accuracy ceiling is entirely dictated by the effective information content in the data (the Shannon coding limit, to be precise).
To ensure uniform data standards from the source, Cai says he spends every day reexamining the definitions and classification systems of data concepts — that is, the ontology — and even repeatedly auditing historical labels. Whether it's setting rules for reinforcement learning or fine-tuning precise annotations, what ultimately determines AI performance is data quality and evaluation rigor, not the model architecture itself.