Shot in Tamarindo, Costa Rica
Sicong Jiang
Building RSI Agents @ Google DeepMind | PhD @ McGill University
I am an Incoming Research Scientist at
Google DeepMind, where I pioneer Agentic Recursive Self-Improvement (RSI). Previously, as a Founding Scientist at
Abaka AI, I directed evaluation research and architected large-scale agentic data & RL environment systems, delivering mission-critical datasets and infrastructure to several frontier AI labs.
My research focuses on the fundamental challenge of building self-evolving AI agents through the lens of automated evaluation, reward modeling, and verifier harnesses. I develop high-impact benchmarks and alignment suites across the agent stack: from sandboxed evaluation harnesses (Harbor-Index, VeriWeb) and human-aligned reward/world models (EditReward, WorldReasonBench) to tool-augmented multimodal reasoning (AgentThink, ChartNet). My mission is to engineer reliable intelligence capable of open-ended, autonomous self-evolution.
Aug 2026
๐ Excited to join
Google DeepMind as a Research Scientist in Oct 2026 to keep working on Gemini RSI!
Aug 2026
๐ One paper accepted by WACV 2027. Check EvaDrive.Jul 2026
๐ Released Harbor-Index โ one of the most challenging agentic benchmarks to date. Great teamwork!Jun 2026
๐ ChartNet featured in MIT News! Grateful for the fantastic collaboration with the MIT-IBM Lab.Mar 2026
๐ Excited to join
Google DeepMind (London) as a Research Intern to work on Self-Evolving Agents.
Jan 2026
๐ One paper accepted by ICLR 2026. Check EditReward.Nov 2025
๐ One paper accepted (oral) by Bridge Program of AAAI 2026.Aug 2025
๐ One paper accepted by EMNLP 2025. Check AgentThink.Jul 2025
๐ One paper accepted by ICCV 2025 Foundation Models for AD Workshop. Check VLA4AD Survey.Mar 2025
โ๏ธ Invited to contribute to Humanity's Last Exam, an AGI reasoning benchmark.Feb 2025
๐ One paper accepted by ICLR 2025 Trustworthy LLM Workshop. Check SparseAttack-LLM4TS.Jan 2025
๐ One paper accepted by AISTATS 2025. Check Attack-LLM4TS.* indicates equal contribution. For full list, visit Google Scholar.
AI Agents, Benchmarks & Evaluation
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
K. Wu*, S. Jiang*, M. Ku, P. Nie, M. Liu, W. Chen
ICLR 2026
Website • Paper • GitHub โญ 158 • HuggingFace ๐ฅ 2K Downloads
AgentThink: Tool-Augmented Reasoning in VLMs for Autonomous Driving
K. Qian*, S. Jiang*, Y. Zhong*, Z. Luo, Z. Huang, et al.
EMNLP 2025
Website • Paper • GitHub โญ 147
VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking
2077AI Team
Under review, 2025
Website • Paper • GitHub โญ 88 • HuggingFace ๐ฅ 5K Downloads
Harbor-Index: A Compact, Diverse, Challenging Benchmark for Agentic Evaluation
Terminal Bench Team
Technical Report & Benchmark Release, 2026
Website • Harbor Hub • GitHub โญ 4.4K
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
X. Gao, S. Jiang, B. Liu, X. Chen, M. Yang, S. Yang, M. Wu, et al.
Under review, 2026
Website • Paper • GitHub โญ 188
WorldReasonBench: Stress-Testing Video Generators as World-State Predictors
K. Wu*, Y. Cui*, W. Xue, Q. Wang, X. Luo, Z. Feng, Z. Yang, S. Wang, S. Jiang, et al.
Under review, 2026
Website • Paper • GitHub โญ 23 • HuggingFace ๐ฅ 2K Downloads
EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks
L. Liu, D. Li, Y. Liang, S. Jiang, H. Vijay, H. Hu, et al.
CVPR 2026 Findings
Website • Paper
Foundation Models: Robustness, Safety & Applications
A Survey on VisionโLanguageโAction Models for Autonomous Driving
S. Jiang*, Z. Huang*, K. Qian*, Z. Luo, T. Zhu, et al.
ICCV Workshop, 2025
Paper • GitHub โญ 613 • Tech Channel Report
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting
F. Liu*, S. Jiang*, L. Miranda-Moreno, S. Choi, L. Sun
AISTATS 2025
Paper • GitHub โญ 17
EvaDrive: Evolutionary Adversarial Policy Optimization for End-to-End Autonomous Driving
S. Jiao, K. Qian, H. Ye, Y. Zhong, Z. Luo, S. Jiang, Z. Huang, et al.
WACV 2027
Paper
FASIONAD+: Enhanced Safety in Autonomous Driving with Adaptive Feedback
Z. Luo*, S. Jiang*, K. Qian*, Z. Huang, J. Miao, et al.
ICRA 2026
Paper
MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving in Corner Cases
Z. Luo*, K. Qian*, J. Wang, Y. Luo, J. Miao, Z. Fu, Y. Wang, S. Jiang, Z. Huang, et al.
ICRA 2026
Paper
Communication-Aware Reinforcement Learning for Cooperative Adaptive Cruise Control
S. Jiang, S. Choi, L. Sun
TRB Annual Meeting (Oral), 2024
Paper
Self-Evolving Agents: Engineered an execution-grounded RL architecture for autonomous self-improvement, enabling Gemini Flash-tier models to outperform Pro-tier baselines on competitive coding benchmarks.
Test-Time Compute Distillation: Designed an execution-consistency verification pipeline, internalizing test-time scaling compute into permanent policy weight updates through continuous RL.
Harness-Data Co-Evolution: Architected an automated harness pipeline to actively harvest high-hardness RL data and edge cases, driving continuous model gains via the co-evolution of evaluation harnesses and trajectory data.
RL Diagnostics & Reward Hacking: Developed an automated pipeline to identify and mitigate reward hacking loops; enhanced training stability and policy performance by filtering anomalous agentic rollouts in the RL loop.
Research: As a founding member of the Research team, I lead benchmarking and evaluation for agentic and multimodal LLMs. I led the EditReward (ICLR'26) project and co-developed large-scale benchmarks including SuperGPQA (NeurIPS'25), ChartNet (CVPR'26), EgoTL (CVPR'26) and VeriWeb.
Advanced Dataset & Pipeline Design: Led several zero-to-one pipeline buildsโarchitecting and deploying high-difficulty dataset solutions and production pipelines from scratch across coding, IMO-level math, multimodal data, agentic trajectories, and RL environments. These datasets and pipelines are directly used for model training and evaluation for several frontier AI labs.
As a core contributor, conducting substantial research across benchmarks, datasets, and agent evaluation for the open-source community.
Agent Evaluation: Led research on agent evaluation and training datasets, focusing on long-horizon reasoning, tool use, and self-evolving agent capabilities.
Multimodal Image Datasets: Led multimodal dataset research for image generation, including preference data and evaluation frameworks for alignment and controllability.
Multimodal Data Pipelines: Built data pipelines and multi-stage QA systems for multimodal LLM projects, overseeing large-scale annotation workflows and label consistency.
Dataset Quality & Validation: Conducted analysis and validation to refine annotations and ensure robust datasets for LLM post-training.
AgentThink (Agent Reasoning): Led a collaboration with Xiaomi and Tsinghua on tool-augmented reasoning for vision-language models in autonomous driving, achieving +54% answer accuracy on open-source models.
Adversarial LLM4TS: Developed a black-box attack framework and public benchmarks for LLM-based time-series forecasting, in collaboration with the Amazon Chronos and Nixtla teams.
Multi-Agent RL Exploration: Developed a multi-agent search strategy combining MADDPG with frontier-based exploration, and built evaluation benchmarks for exploration efficiency.
Awards
2024
McGill Engineering Doctoral Award (MEDA)2021
TISED Doctoral Recruitment Award (DRA), McGill University2019
Outstanding Graduate of Liaoning Province; Most Influential Graduate, Northeastern University2017
National 1st Prize, China Undergraduate Mathematical Contest in Modeling2017
1st Class Academic Scholarship, Northeastern UniversityAcademic Service
Workshops Organizer
- CVPR 2026 Workshop on Video Generative Models: Benchmarks and Evaluation
- ICCV 2025 Workshop on Memory and Vision
- COLM 2025 Workshop on AI Agents: Capabilities and Safety
Conferences Reviewer
- Advances in Neural Information Processing Systems (NeurIPS)
- International Conference on Learning Representations (ICLR)
- International Conference on Artificial Intelligence and Statistics (AISTATS)
- IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- International Conference on Computer Vision (ICCV)
- Conference on Language Modeling (COLM)
- Conference on Empirical Methods in Natural Language Processing (EMNLP)
- Association for the Advancement of Artificial Intelligence (AAAI)
- IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
- IEEE International Conference on Robotics and Automation (ICRA)
- IEEE Intelligent Transportation Systems Conference (ITSC)
Journals Reviewer
- IEEE Robotics and Automation Letters (RA-L)
- Transportation Research Part C: Emerging Technologies (TRC)
- IEEE Transactions on Intelligent Transportation Systems (T-ITS)
I enjoy music by Tyler, the Creator, SZA and Chappell Roan.
Sometimes I also listen to Taylor Swift, Olivia Rodrigo and 9m88.
My favorite influencer is Allywoo on RedNote.
Cat: Bobo, a golden shaded British Shorthair who is good at programming with buttons.




