About me

I research and build reliable AI agents that can reason, learn from feedback, and turn research ideas into systems that work in the real world.

At TaoTian Group @ Alibaba, my current work focuses on AI agent research across three core directions: AutoResearch, post-training, and agentic reinforcement learning. I apply these ideas in Xianyu AI systems and am especially interested in closed-loop research agents, stable reasoning objectives, and high-quality data systems for foundation models.

Before Alibaba, I contributed to foundation-model development at Tencent and Baidu. At Tencent, I worked on the Hunyuan Foundation Model, focusing on mathematical and biomedical capability enhancement, pre-training data, multilingual capability improvement, and Yuanbao AI Search. At Baidu, I contributed to multilingual capability enhancement for the ERNIE Bot 5 (EB5) Foundation Model.

Earlier, at the Chinese Academy of Sciences, I conducted research on knowledge graphs and LLM-based security.

I am always happy to connect and collaborate on ambitious research problems. Feel free to reach out by email or visit my Google Scholar.

Current research

I study how agents can research, learn, and act reliably across three core directions, with Xianyu AI as a practical application domain.

AutoResearch

Closed-loop agents for scientific discovery.

Post-Training

Data and alignment for stronger reasoning.

Agentic RL

Reliable reinforcement learning for tool-using agents.

In progress
CVPR Manuscript Agent Research Survey

Publications & Manuscripts

SILICA shared-state counterfactual identifiability evaluation

Submitted

SILICA: Certified Counterfactual Evaluation of Identifiability in Unseen-Language Induction

Submitted to EACL · Aug 2026 · Not yet public

Pengcheng Xu, Xinyu Guan

Evaluates whether an unseen-language task is identifiable by constructing same-state counterfactual queries with provably different correct actions.

In Preparation

When KL Regularization Fails in Online Reasoning RL: A Token-Level Gradient Contract

Withdrawn from AAAI · In preparation for ICLR · Aug 2026 · Not yet public

Dingding, Runhao Liu, Yongkang Zhang, Zijian Zeng, Yuhao Liao, Xinyu Guan, Huiming Yang

Frames KL failures as a token-level gradient-contract problem and proposes ZCPO to enforce zero-sum shared-prefix gradients.

MaxNorm-AC advantage-scale calibration pipeline

Submitted

Advantage Scale Calibration Imbalance in Group-Relative Optimization under Low-Variance Rewards: Diagnosis and Bounded Recovery

Submitted to AAAI 2027 · Jul 2026 · Not yet public

Dingding, Runhao Liu, Yongkang Zhang, Zijian Zeng, Yuhao Liao, Xinyu Guan, Huiming Yang

Diagnoses calibration imbalance under low-variance rewards and introduces bounded recovery for credible small reward gaps.

ChronoMem event-memory and residual-forecasting pipeline

In Preparation

ChronoMem: Interpretable Event Memory for LLM-Augmented Time-Series Forecasting

In preparation for ICASSP 2027 · Aug 2026 · Not yet public

Xinyu Guan et al.

Uses LLMs to structure external events as an interpretable residual memory rather than directly forecasting numerical values.

Patents

  • Under ReviewA Multi-Agent and LLM Collaborative Multi-Dimensional Text Quality Scoring System
  • GrantedA Generative Large Model Watermarking Tool Based on Probability Perturbation Encryption
  • GrantedA Database Drag Behavior Detection Method Based on Time Series

Work experience

Feb 2026 — Present

TaoTian Group @ Alibaba

AI Agent Researcher · P6

AI Agent research spanning General AutoResearch and multimodal quality inspection for Xianyu.

  • General AutoResearch — Re-architected the Agent runtime with bounded loop control, stalled-branch rerouting, evidence-driven hypothesis generation, hierarchical memory, and resumable execution traces. The system autonomously performs bad-case diagnosis, candidate generation, and prompt or policy iteration. Across 19 business-task optimization runs, it completed prompt iteration, candidate validation, and migration validation without human intervention within the optimization loop. Across 4 Agent runtimes × 4 model configurations, the selected stack met target on 8 of 9 controlled image-QC tasks, with 279 migration tests passed.
  • Xianyu Multimodal Quality Inspection — Built inspection workflows for viewpoint compliance detection, base photo-quality checks, and visible physical-defect detection. An offline acceptance snapshot limited to nine categories and 80 inspection checks reached 97.54% mean Macro-F1. Separately, the current capability supports 30+ product categories at approximately 30K orders per day.

Oct — Dec 2025

Baidu / ERNIE Foundation Model Core Team

Senior Research Scientist · T4+

DAPO-based post-training and alignment for EB5 on PaddlePaddle.

  • Scaled multilingual post-training to 260 H800 GPUs across 80+ languages; reached a Text Arena win rate of 86.4% and multilingual average accuracy of 92%, tied global #2 and China #1 on LMArena in Nov 2025.
  • Led ten-axis instruction evolution and a best-of self-play plus human N-to-1 data pipeline; achieved a +25.5% win-rate result against EB45T-VL-SFT-v13 and tied #1 on the internal Vision-Arena.

Mar 2025 — Sep 2025

Tencent / Hunyuan Text-to-Text Pipeline Team

Research Scientist · T5

Yuanbao AI Search with NeuroBERT-pro and GATv2Conv.

  • Trained on 8 H100 GPUs and deployed on 800 H100 GPUs for about 50B HTML pages/day; extraction reached 65%, classification improved 70% → 80%, and structural accuracy improved 60% → 65%.
  • Improved maximum-rectangle precision 83% → 98% and recall +1.5%; HTML pruning raised recall 73.41% → 98.33% with 57.14% compression and reduced pretraining tokens by over 50%.
  • Produced approximately 900B multilingual tokens through instruction evolution and data synthesis.
  • Developed and maintained STEM main-text extraction operators, completing 135 deliveries and ranking 2nd among 35 contributors in the closed research program.

Feb 2024 — Mar 2025

Tencent / Hunyuan Strategy Group 4

Research Scientist · T5

Mathematical and biomedical capability enhancement, video processing, data recognition, and audio alignment.

  • Built a GPT-4o-based agent pipeline with 99% concept-extraction precision and curated 934 high-quality questions; MMLU improved +11.7 and CMMLU improved +10.2.
  • Processed a 10-minute 1080p video in about 80 seconds, with capacity of 20,000 video-hours per month.
  • Magika exceeded 99% across 116 file types, running at 10 QPS for 800,000 files/day.
  • Whisper alignment reached 99% and delivered 20,000+ transcripts, contributing to a patent.

Nov 2023 — Feb 2024

Institute of Information Engineering, Chinese Academy of Sciences

Research Assistant Intern

Fine-tuned LLM approaches for insider-threat detection and structured knowledge-graph construction.

  • Worked with OPT-125M, ChatGLM-6B P-tuning, and LLaMA-7B LoRA on CERT 4.2; AUC and Recall were about 0.99, FPR about 0.3%, training time 30% lower, and compute over 50% lower.
  • Curated 20,000+ RDF triples with 98% extraction accuracy and precision +43%.

May 2021 — May 2022

Mico World / Yoho Department

Software Engineer

Frontend architecture and production systems for international social products.

  • Greedy Lion H5 generated about 20% of company revenue; critical-path optimization improved load speed 3.7×.
  • WakaDashboard delivered over 80% of core APIs and raised engineering efficiency 73%.

Research experience

Nov 2023 — Feb 2024

LLM4ITD: Insider Threat Detection with Fine-Tuned LLMs

Reformulated classification tasks as question-answering with structured prompts and parameter-efficient tuning of about 0.01% of parameters on CERT 4.2. AUC reached 0.9923, Recall 98.83%, training time was 30% lower, and compute was over 50% lower.

Aug 2023 — Oct 2023

Yellow Crane Tower Tourism Dialogue System

Extended and reparameterized RoPE and applied LoRA to Llama 2 7B. Built approximately 15,000 GPT-4-distilled tourism dialogues, with about 89% answer accuracy in human evaluation.

Jun 2023 — Oct 2023

Efficient Text Search Algorithm Evaluation

Collaboration with Prof. David Manlove. Evaluated seven algorithms on a 1M+ NLTK corpus and developed an Ukkonen-based suffix-tree approach with approximately 5.2× speedup. Applied it to human genomic sequences, improving search speed by 40% and reaching 100% accuracy in the reported evaluation.

Apr 2023 — Jul 2023

Lung Cancer Literature Classification with BioBERT

Expanded the dataset by approximately 6,000 samples through synonym replacement and back-translation, fine-tuned BioBERT with hierarchical freezing, and reached 93% PubMed accuracy with about 15% improvement on the manual evaluation set.

Jan 2023 — Jun 2023

EEG Feature Analysis for SCI Patients

Applied feature-selection methods to 5,000 samples: KNN improved 69.4% → 77.8%, and SVM reached 94.4%.

Nov 2022 — Jan 2023

ML Analysis of WSI Colorectal Cancer Datasets

Applied BERT-based augmentation to around 25% of the dataset and evaluated KMeans, Louvain, PCA, and UMAP.

Education

Sep 2022 — Dec 2023

University of Glasgow

MSc in Computer Science

GPA 3.67 / 4.0

Research collaboration with Prof. David Manlove, University of Glasgow professor and University of Oxford graduate.

Sep 2016 — Jun 2020

Hubei University

BEng in Software Engineering