A Review of Automated Data Science Based on Large Language Models
A Review of Automated Data Science Based on Large Language Models
Weixiong Cui,Gulan Zhang,Baoli Wang,Xu Su
Abstract
Large Language Models (LLMs) are revolutionizing automated data science by enabling adaptive, natural languagedriven workflows that transcend the limitations of traditional AutoML tools. This review comprehensively examines how LLMagent systems automate end-to-end data science pipelines. We trace the evolution from manual methodologies and rule-based AutoML solutions toward dynamic frameworks leveraging LLM capabilities. Three core architectural paradigms are analyzed: single-agent systems for linear orchestration, function-oriented multi-agent systems with rigid role specialization, and roleoriented multi-agent systems supporting dynamic task transfer. Critical challenges include semantic misalignment in tool calls, limited cross-modal reasoning, and evaluation gaps in real-time adaptability. Future research must prioritize unified modeling for heterogeneous data, reinforcement learning-driven agent collaboration, and trustworthy AI mechanisms incorporating causal interpretability and domain-specific fairness constraints.
