Jiayu Song.

Learning · Perception · Robotics

Jiayu Song

UT Dallas monogramPh.D. student at UT Dallas

Teaching machines to connect the dots—between what they see, what they understand, and what they do. Exploring multimodal learning, reinforcement learning, and robotics along the way.

Jiayu Song’s illustrated avatar
The University of Texas at Dallas

01 / ABOUTA little about me

Hi, I’m Jiayu! I’m a Ph.D. student at UT Dallas, advised by Prof. Feng Chen. I’m drawn to a deceptively simple question: how can machines make sense of the world and learn what to do next? That curiosity connects my interests in multimodal learning, reinforcement learning, and robotics.

Previously, I received my B.S. (2021) and M.S. (2024) from the School of Computer Science and Engineering at Central South University Central South University emblem, where I worked in the Multimedia Lab under the supervision of Prof. Shichao Zhang.

If you’re thinking about similar questions, I’d love to compare notes—say hello!

  • Multimodal learning
  • Reinforcement learning
  • Robotics

02 / EDUCATIONAlong the way

  1. Current
    UT Dallas monogram

    The University of Texas at Dallas

    Ph.D. student · Advisor: Prof. Feng Chen

  2. Central South University emblem

    Central South University

    M.S. · School of Computer Science and Engineering

  3. Central South University emblem

    Central South University

    B.S. · School of Computer Science and Engineering

03 / RESEARCHPublications

Four selected papers. Find the full list on Google Scholar ↗.

  1. CoRDE framework: concept priors guide routing to LoRA diffusion experts for robot manipulation

    IEEE/RSJ IROSAccepted

    CoRDE: Concept-Prior Routed Diffusion Experts for Structural Generalization in Robot Manipulation

    Haidong Huang, Xixin Zhao, Yaohua Zhou, Jiayu Song, Jiayi Zhang, Jun Ma, Haiyue Zhu, Xiaocong Li

  2. MSSPQ framework with image and text embeddings, hashing layers, and multiple correlation losses

    ACM ICMR

    MSSPQ: Multiple Semantic Structure-Preserving Quantization for Cross-Modal Retrieval

    Lei Zhu, Liewu Cai, Jiayu Song, Xinghui Zhu, Chengyuan Zhang, Shichao Zhang

  3. MG-HSF framework: cross-modal feature embeddings, multi-graph hierarchical semantic fusion, and adversarial learning

    IEEE ICME

    Multi-graph Based Hierarchical Semantic Fusion for Cross-modal Representation

    Lei Zhu, Chengyuan Zhang, Jiayu Song, Liangchen Liu, Shichao Zhang, Yangding Li

  4. HCMSL framework combining cross-modal representation learning with image and text intra-modal similarity learning

    ACM TOMM

    HCMSL: Hybrid Cross-Modal Similarity Learning for Cross-Modal Retrieval

    Chengyuan Zhang, Jiayu Song, Xiaofeng Zhu, Lei Zhu, Shichao Zhang