Ben Yu

yubenjamin2022 at ucla dot edu

Hi! I'm Ben, an undergraduate at UCLA, majoring in Data Theory. I am currently a researcher with the PLUS Lab of the UCLA NLP Group, advised by Wenbo Hu and Professor Violet Peng. Here, I am researching multimodal & unified models, exploring how we can use generation objectives to improve their reasoning and understanding capabilities. I was formerly a researcher with the Digital Synthesis Lab, advised by Professor Daniel Schwalbe-Koda. There, I worked on information-theoretic approaches to efficiently compress atomistic datasets. I was also previously a researcher with the Computational Machine Learning Group, advised by Justin Cui. There, I researched improving performance of image and video generation models in text preservation and efficient generation.

At a philosophical level, I want to understand what it means for humans to reason and understand the world. I believe that humans think through internal representations independent of our spoken language, an idea which motivates my interest in developing AI systems which compute in language-agnostic representations.

Through my work, I hope to explore: (i) grounded approaches that enable AI systems to reason and perceive in more human-like ways, and (ii) evaluation methods that can effectively measure these capabilities. In the end, I hope to gain insight into how humans learn and reason by studying these processes through the lens of artificial systems. I'm always looking for new and interesting research opportunities in these areas, so feel free to reach out if you'd like to collaborate!

Email  /  CV  /  Scholar  /  GitHub

profile photo

Publications

* Represents Equal Contribution
Smart-GRPO: Smartly Sampling Noise for Efficient RL of Flow-Matching Models
Benjamin Yu, Jackie Liu*, Justin Cui
AAAI 2026 AIR-FM Workshop
arXiv

Using optimized noise perturbations for reinforcement learning, we enable flow-matching models to efficiently improve image quality and human alignment.

Maximizing Efficiency of Dataset Compression for Machine Learning Potentials With Information Theory
Benjamin Yu, Vincenzo Lordi, Daniel Schwalbe-Koda
Journal of Chemical Physics
arXiv, GitHub

By quantifying dataset redundancy using information entropy, we can selectively subsample data to achieve compact, information-preserving datasets that improve training efficiency.

Miscellanea

Oral Presentations

Efficient compression of atomistic datasets with information theory, APS Global Physics Summit 2025

Academic Service

Not yet but hopefully eventually!

Teaching

ACM AI at UCLA, Workshop Officer, 2025 - 2026
Statistics Club Workshop Chair, 2024 - 2025

Visitor Map


Thank you to Jon Barron for this website template, the source code is here.