My research focuses on building AI systems that are explainable, robust, and scalable. I build agentic AI systems that reason, plan, and collaborate across long-horizon tasks — designing multi-agent architectures with composable reasoning modules, grammar-based coordination protocols, and scalable frameworks that let autonomous agents tackle complex, multi-step problems. My work also focuses on creating rigorous benchmarks to systematically evaluate LLMs and VLMs, probing their capabilities and surfacing critical failure modes.
I am especially interested in specialized agentic systems for vertical AI domains — bringing long-horizon reasoning and multi-agent collaboration to industry-specific challenges where domain expertise, reliability, and safety are paramount.
My recent research has spanned the full LLM and VLM lifecycle — continual pre-training, post-training (DPO variants, safety, abstention), and efficient inference (speculative decoding, dynamic layer slicing, layer-wise quantization) — actively publishing at top-tier conferences including ACL, EMNLP, ICLR, NAACL, AAAI, SIGIR, and COLM. I received my Ph.D. from the University of Arizona and have spent over 6 years in industry research, building and deploying large-scale AI systems across diverse real-world applications.
Get in TouchMy work spans the full LLM lifecycle — from continual pre-training to post-training alignment — with a focus on building scalable, efficient, and aligned AI systems.
Long-horizon reasoning, chain-of-thought optimization, and multi-step planning systems. Developing benchmarks for evaluating and mitigating over-reasoning in LLMs.
Autonomous, scalable agentic systems for complex multi-environment tasks. Grammar-based search and composable reasoning modules for generalizable agent collaboration.
Speculative decoding, dynamic layer slicing, and layer-wise quantization techniques for accelerating LLM inference while preserving quality.
Preference alignment via DPO variants, adversarial robustness of reasoning models, abstention capabilities, and defense against prompt injection attacks.
Enhancing LLM capabilities across non-Latin scripts through phonemic prompting, multilingual instruction alignment, and cross-lingual transfer learning.
Chart reasoning with visual reinforcement, RAG-powered document QA, and multi-modal dialogue systems integrating vision and language understanding.
20 most recent papers — sorted by recency. Click any card for details.
Or email directly — vikasy.zona@gmail.com