AutoResearch
Closed-loop agents for scientific discovery.
I research and build reliable AI agents that can reason, learn from feedback, and turn research ideas into systems that work in the real world.
At TaoTian Group @ Alibaba, my current work focuses on AI agent research across three core directions: AutoResearch, post-training, and agentic reinforcement learning. I apply these ideas in Xianyu AI systems and am especially interested in closed-loop research agents, stable reasoning objectives, and high-quality data systems for foundation models.
Before Alibaba, I contributed to foundation-model development at Tencent and Baidu. At Tencent, I worked on the Hunyuan Foundation Model, focusing on mathematical and biomedical capability enhancement, pre-training data, multilingual capability improvement, and Yuanbao AI Search. At Baidu, I contributed to multilingual capability enhancement for the ERNIE Bot 5 (EB5) Foundation Model.
Earlier, at the Chinese Academy of Sciences, I conducted research on knowledge graphs and LLM-based security.
I am always happy to connect and collaborate on ambitious research problems. Feel free to reach out by email or visit my Google Scholar.
I study how agents can research, learn, and act reliably across three core directions, with Xianyu AI as a practical application domain.
Closed-loop agents for scientific discovery.
Data and alignment for stronger reasoning.
Reliable reinforcement learning for tool-using agents.
In Preparation
Studies when multimodal agents should refine prompts or acquire additional visual evidence, with feedback-driven intervention selection and independent verification.
Submitted
Combines source-aware evidence organization, state-conditioned historical response transfer, and reliability-guided fusion for event-informed time-series forecasting. Code
Submitted
Learns when a frozen vision-language-action model should exit early or keep computing, using budget-constrained RL while the simulated environment continues to evolve.
Submitted
Isolates a parameter-pool confound in LoRA masking: density-matched random B-matrix masks reduce loss-based forgetting across seven seeds, while accuracy evidence remains inconclusive.
Submitted
Builds static RAG chunk indices from question–evidence supervision, unifying overlapping and non-overlapping chunking as budget-aware utility maximization.
Submitted
Evaluates whether an unseen-language task is identifiable by constructing same-state counterfactual queries with provably different correct actions.
In Preparation
Frames KL failures as a token-level gradient-contract problem and proposes ZCPO to enforce zero-sum shared-prefix gradients.
Submitted
Diagnoses calibration imbalance under low-variance rewards and introduces bounded recovery for credible small reward gaps.
Accepted
Introduces CICL, a decision-aware context layer that selects evidence by its expected effect on an agent's next action.
Preprint
Presents a pattern-matching approach based on Ukkonen's construction for efficient text search.
Published
Introduces a basket-enhanced heterogeneous hypergraph for price-sensitive next-basket recommendation. arXiv
Feb 2026 — Present
AI Agent Researcher · P6
AI Agent research spanning General AutoResearch and multimodal quality inspection for Xianyu.
Oct — Dec 2025
Senior Research Scientist · T4+
DAPO-based post-training and alignment for EB5 on PaddlePaddle.
Mar 2025 — Sep 2025
Research Scientist · T5
Yuanbao AI Search with NeuroBERT-pro and GATv2Conv.
Feb 2024 — Mar 2025
Research Scientist · T5
Mathematical and biomedical capability enhancement, video processing, data recognition, and audio alignment.
Nov 2023 — Feb 2024
Research Assistant Intern
Fine-tuned LLM approaches for insider-threat detection and structured knowledge-graph construction.
May 2021 — May 2022
Software Engineer
Frontend architecture and production systems for international social products.
Nov 2023 — Feb 2024
Reformulated classification tasks as question-answering with structured prompts and parameter-efficient tuning of about 0.01% of parameters on CERT 4.2. AUC reached 0.9923, Recall 98.83%, training time was 30% lower, and compute was over 50% lower.
Aug 2023 — Oct 2023
Extended and reparameterized RoPE and applied LoRA to Llama 2 7B. Built approximately 15,000 GPT-4-distilled tourism dialogues, with about 89% answer accuracy in human evaluation.
Jun 2023 — Oct 2023
Collaboration with Prof. David Manlove. Evaluated seven algorithms on a 1M+ NLTK corpus and developed an Ukkonen-based suffix-tree approach with approximately 5.2× speedup. Applied it to human genomic sequences, improving search speed by 40% and reaching 100% accuracy in the reported evaluation.
Apr 2023 — Jul 2023
Expanded the dataset by approximately 6,000 samples through synonym replacement and back-translation, fine-tuned BioBERT with hierarchical freezing, and reached 93% PubMed accuracy with about 15% improvement on the manual evaluation set.
Jan 2023 — Jun 2023
Applied feature-selection methods to 5,000 samples: KNN improved 69.4% → 77.8%, and SVM reached 94.4%.
Nov 2022 — Jan 2023
Applied BERT-based augmentation to around 25% of the dataset and evaluated KMeans, Louvain, PCA, and UMAP.
Sep 2022 — Dec 2023
MSc in Computer Science
GPA 3.67 / 4.0
Research collaboration with Prof. David Manlove, University of Glasgow professor and University of Oxford graduate.
Sep 2016 — Jun 2020
BEng in Software Engineering