Post-training
Reinforcement learning and verifiable feedback for language, scientific, and agentic systems.
system.profile / online
I develop methods for large language models, reinforcement learning, and multimodal scientific foundation models. I am a Computer Science PhD student at UC Irvine, with an emphasis on reliable evaluation.
Curated commands only—this terminal does not execute
code. Try help.
trace://research-loopillustrative · not live telemetry
Reinforcement learning and verifiable feedback for language, scientific, and agentic systems.
Turning large general models into smaller, efficient experts without losing what matters.
Foundation models and evaluations grounded in genomics, biology, diagrams, and geometry.
Quick route
stream.01 / latest
Accepted work, research roles, and recent milestones—without the notification noise.
LLM Research Scientist Intern working on verifiable evaluation and post-training for scientific diagrams.
Differentiable combinatorial optimization for causal variant discovery in the non-coding genome.
Procedurally generated tasks for long-horizon video reasoning.
Section-wise retrieval, LLM generation, and reinforcement learning for biomedical lay summaries.
For work on L2 normalization and geodesic distance in high-dimensional single-cell visualization.
workspace.02 / current
Methods and systems work, organized by the scientific question rather than the application domain.
R/01
Project lead · model distillation, evaluation, and systems
Toward sub-million-parameter expert models distilled from large genomic language models, with controlled evaluation across classification and base-resolution tasks.
R/02
Project lead · model, data, training, and evaluation
A reasoning-grounded, multimodal foundation model for cell-type-conditioned DNA generation, editing, and pair prediction. I lead the model, data, and evaluation effort.
R/03
LLM Research Scientist Intern · diagram.ai
Semantic verifiers turn structured diagram feedback into rewards for reinforcement-learning post-training and iterative repair. The public materials emphasize evaluation design and reproducible experiments; unreleased results remain high level.
R/04
Independent study · controlled empirical research
Controlled multi-domain studies across NLP and single-cell models, with multi-seed baselines, negative controls, and reproducible cluster-scale experiments.
R/05
LLM agents · knowledge distillation · on-policy training
Studying stable knowledge transfer across multi-turn agent trajectories, supported by multi-GPU training and evaluation infrastructure built with FSDP, vLLM, Ray, and SLURM.
R/06
Technical report · distributed ML systems
Hardware-aware, latency-predictable differentiable search for faster configuration and convergence of distributed machine-learning pipeline parallelism.
R/07
Public preprint · interpretable computational biology
Interpretable multi-task learning for shared epigenetic regulation across autoimmune diseases, with site–gene–pathway structure built into the analysis.
R/08
Open benchmark contributions · Kart and Minecraft proposals
I designed long-horizon video-understanding proposals for race-telemetry reconstruction in SuperTuxKart and action-ledger reconstruction in Minecraft, with seeded generators, machine-exact ground truth, deterministic graders, and anti-shortcut calibration.
Kart issue #73 ↗ Minecraft issue #74 ↗ Kart branch ↗ Minecraft branch ↗
method.lab / transfer
Two inspectable method notes: the manuscript-derived OmegaGenome loss anatomy and established on-policy curricula for multi-turn agents. Unreleased results are intentionally omitted.
archive.03 / selected
Selected peer-reviewed work across genomic discovery, video reasoning, language models, and single-cell geometry.
Efficient differentiable search for causal variants across molecular modalities and tissues.
Read paper
Section-wise retrieval, LLM generation, and reinforcement learning for accessible biomedical summaries.
Read paper
Project cells to the L2 hypersphere, compare them by angular distance, then use a spherical affinity inside SNE.
Transformer-based reinforcement learning for oracle-guided molecular de novo design.
Read paperAdapts language models for accessible biomedical communication, emphasizing readability and factual quality.
paper.controls / inspect
Move a published optimization parameter and inspect an exact reported ablation. Illustrative quantities remain separate from experimental measurements.
EARLIER / PUBLIC RECORD
Author PDFs, full author lists, and figure notes Open the academic publication section →
paper.lab / interactive reconstruction
An interactive reading of our ACM BCB 2024 paper. Move κ to see how fixed angular distances become a sharper or broader probability neighborhood.
ACM BCB 2024 · ACM SIGBio Best Paper Award
The method projects each cell to the unit hypersphere, measures angular distance, transforms angles into spherical affinities, and normalizes those affinities before optimizing the low-dimensional embedding. The interaction below makes that causal chain inspectable.
Publication ↗ Author PDF ↗ Method figure ↗ Embedding figure ↗ Evaluation figure ↗
The geometry and κ controls stay in view through this method trace; follow the numbered formulas from projection to the low-dimensional neighborhood.
systems.04 / field log
I move between objectives, datasets, distributed systems, product interfaces, and evaluation infrastructure.
research / post-training
Verifiable evaluation and reinforcement-learning post-training for mathematical and scientific diagram generation.
research / scientific ML
Antibody binding-affinity models, leakage-resistant splits, balanced objectives, and reproducible evaluation.
product systems
React/TypeScript workflows, Java/AWS backend integrations, validation, and anomaly detection for VMware Cloud on AWS.
open-source ML systems
Built an automated feature-generation and selection workflow with Python and OpenMLDB SQL, contributed it upstream, and presented it to the open-source community.
Contribution ↗ 4Paradigm ↗ OpenMLDB ↗ Meetup talk ↗ Code Camp talk ↗multimodal learning
Worked on multimodal target detection with zero-shot depth estimation and multimodal neural architecture search.
Shanghai AI Laboratory ↗model efficiency / open source
Contributed model-compression and quantization work and documentation to Intel Neural Compressor; studied inference-server architecture and helped implement C++ inference tooling.
Intel Neural Compressor ↗distributed medical AI
Built multi-node, multi-GPU 3D U-Net training with Horovod, OpenMPI/NCCL, and NVIDIA Clara, and added distributed-training backend support for medical-imaging workloads.
Project repository ↗ Shukun Technology ↗ Horovod ↗ Technical talk ↗working stack
TEACHING
Teaching Assistant for UCI ICS 6B and ICS 6D, and University of Michigan–Shanghai Jiao Tong University Joint Institute VE370.
SELECTED HONORS
ACM SIGBio Best Paper Award · SJTU Outstanding Graduate · Microsoft Imagine Cup China third prize · Mathematical Contest in Modeling Meritorious Winner.
MCM paper ↗SYSTEMS MODE
Python, C/C++, Java, TypeScript, SQL, shell, CUDA, PyTorch, TensorFlow, Horovod, distributed training, evaluation, and benchmark design.
playground.05 / shipped
Browser worlds, interactive mathematical essays, and tools for thinking with models.
live multiplayer / browser-native
BUILD/01
A solo-built, no-install browser RPG with six game modes, including real-time multiplayer battles—spanning WebGL rendering, Deno services, serverless Postgres, and automated browser testing.
BUILD/02 · FIELD GUIDE
The Geometry of Intelligence
Riemannian and information geometry for AI, made visual.
↗
BUILD/03 · INTERACTIVE ESSAY
Particles → Probability
From interacting particles to the mathematics of Fields Medalist Yu Deng.
↗
BUILD/04 · DEVELOPER TOOL
File → Prompt
A private-by-design browser tool that orders project files and turns them into one
structured prompt. Everything stays local.
↗
BUILD/05 · CONVERSATIONAL WEB
PengchengGPT
A conversational interface to my work, background, and public projects.
↗
inspect source: PengchengGPT repository ↗
side.quest / itch.io
offscreen.06 / human context
Music, speculative writing, browser experiments, and the story encoded in a name.
SELECTED PERSONAL MEDIA
A guitar recording away from the research loop. The embedded player loads only after you choose it, so Bilibili receives no request on the initial page load.
Watch on Bilibili ↗NAME / ORIGIN
“Pengcheng” pairs peng—the giant bird of Chinese mythology—with cheng, a journey. My parents chose it as a hope for an ambitious path and a wide horizon. In English, I pronounce my surname like “Hsu.”
AUDIO / VIDEO
I sing, play guitar, and record occasional technology talks.
WORDS / SPECULATIVE
祂 and Father Sun are two forms of the same science-fiction project.
OFFLINE / INTERESTS
Basketball, tennis, table tennis, swimming, reading, science fiction, and learning how things work. Richard Feynman and Tsung-Dao Lee remain enduring inspirations.
I try to make work that contributes something positive and outlasts the moment in which it was made. The full personal note remains in the academic archive.
channel.07 / open
I welcome conversations about large language models, reinforcement learning, reliable AI evaluation, genomics, distillation, scientific machine learning, and unusually ambitious web experiments.