I study the Self-X of large language models: Recursive Self-Improvement (RSI), Self-Assessment, and Self-Generated Data.
My work centers on the Self-X of LLMs, especially Recursive Self-Improvement (RSI): the ability of a model to evolve on its own, continuously, without human supervision. I'm drawn to this because the traditional recipe of more data and more human supervision is starting to hit a ceiling, while today's LLM agents can already generate data, call tools, and run code on their own. So I believe the next step is for a model to drive its own progress, learning and improving by itself. Around RSI, I also explore Self-Assessment and Self-Generated Data.
I am also interested in RL, Post-Training, RAG, Trustworthy LLMs, and Knowledge Reasoning.
RSI of LLMs. Built Zevo, a self-improving system that autonomously evolves language models toward user-specified objectives.
Reliability of LLMs. Developed a Reasoning-based Bias Detector (RBD) to debias LLM-as-a-judge evaluations.
LLM for Finance. Built Fin-RAG, a RAG-enhanced system with dynamic chunking and hybrid retrieval for financial QA.
RSI of LLMs. Proposed DNPO for LLM self-improvement with synthetic data.
LLM for Healthcare. Built an automated multi-modal pipeline for breast ultrasound report generation.
RAG in LLMs. Developed PRCA, a pluggable adapter improving retrieval-augmented QA with black-box LLMs.