Menu

Categories

Tags

Data quality, not algorithms, is AI's real bottleneck, says Tesla engineer

June 21, 2026 | Source: x | AI | 198 views 0 comments

A Tesla AI senior principal engineer named Cai Yunda has a message for anyone who thinks machine learning is all about training models: you're wrong. According to Cai, the popular assumption that ML projects spend 99% of their time on training is backwards. In reality, only about 2% of effort goes into model parameter training. The rest? 50% is evaluation and testing, 40% is cleaning data, and 8% is system integration.

Cai emphasizes that data cleaning and evaluation fundamentally define the ceiling of what AI can learn. If the raw data has fuzzy definitions or inconsistent labels, you're injecting noise from the start. No amount of algorithmic wizardry or hyperparameter tuning can erase that background noise — the model can't correct its own flawed textbook. The final accuracy is bounded by the effective information content in the data, a hard limit rooted in Shannon's coding theory.

To keep data standards consistent, Cai says he spends his days reexamining the definitions and classification systems behind data concepts — what he calls ontology — and even revisiting historical labels. Whether you're setting rules for reinforcement learning or fine-tuning with precise annotations, the deciding factor is always data quality and evaluation rigor, not the model architecture itself.

Leave a Reply

Your email address will not be published. Required fields are marked *