Papers & projects

Research

I work on reinforcement learning, the reasoning of large language models, and trustworthy AI. Papers are listed newest first.

Papers

  • 2026
    PreprintLLM reasoning

    When Context Changes: Understanding Update Failures in LLMs

    Junyu Guo, Yuchen Fang, Shangding Gu, Costas Spanos, James Demmel, Javad Lavaei

    LLM agents can answer with an old value of a variable even when the updated one is still in context, a failure we call stale binding. We introduce the Controlled In-Context Memory (CICM) benchmark, trace the failure to attention drifting toward old values, and correct most of these errors by redirecting attention without further training.

  • 2026
    PreprintLLM reasoning

    LLMs Should Express Uncertainty Explicitly

    Junyu Guo, Shangding Gu, Ming Jin, Costas Spanos, Javad Lavaei

    Trains language models to state their uncertainty in two ways: a verbalized confidence score after reasoning, and explicit uncertainty markers during reasoning. This reduces overconfident errors and gives downstream systems such as retrieval-augmented generation a signal to act on.

  • 2025
    NeurIPS 2025Safe RL

    Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL

    Junyu Guo, Zhi Zheng, Donghao Ying, Ming Jin, Shangding Gu, Costas Spanos, Javad Lavaei

    DRCORL uses a diffusion model to regularize policy learning in offline safe reinforcement learning, and gradient manipulation to resolve conflicts between reward and safety objectives, while keeping inference fast.

  • 2025
    NeurIPS MATH-AI Workshop 2025LLM reasoning

    Meta Thinker: Thinking What AI Thinks

    Junyu Guo, Shangding Gu, Costas J. Spanos, Javad Lavaei

    Asks whether LLMs can pick the right reasoning style for a problem on their own. After comparing five reasoning paradigms on mathematical, logical and commonsense benchmarks, we propose a meta-thinking prompt algorithm that selects or synthesizes a style from the input, improving both accuracy and token efficiency.

  • 2025
    PreprintLLM reasoning

    StyleBench: Evaluating Thinking Styles in Large Language Models

    Junyu Guo, Shangding Gu, Ming Jin, Costas Spanos, Javad Lavaei

    A benchmark that evaluates five reasoning strategies across diverse tasks and model scales, showing when structured reasoning helps and when it only adds cost or new failure modes.

  • 2023
    PreprintMean field games

    Reinforcement Learning for SBM Graphon Games with Re-Sampling

    Peihan Huo, Oscar Peralta, Junyu Guo, Qiaomin Xie, Andreea Minca

    A reinforcement learning algorithm for multi-population mean field games with graphon structure, with convergence guarantees and experiments on an epidemic model.

Under review

  • Under review at Operations Research

    Optimal Portfolio Allocation in Incomplete Markets

    With Prof. Chenxu Li (Peking University) and Prof. Yiwen Shen (HKUST)

    Decomposes the optimal portfolio policy under CRRA utility and characterizes an investor-specific price of risk for incomplete markets. A backward Monte Carlo algorithm estimates that price of risk and computes the policy, with proven convergence rates and an extension to high dimensions.