Skip to content
View yangpengchengstat's full-sized avatar
  • Purdue University

Block or report yangpengchengstat

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
yangpengchengstat/README.md

Pengcheng Yang

Statistics PhD | Machine Learning | LLM Evaluation & Representation Geometry | Robust AI

I am a Statistics PhD working on machine learning, LLM evaluation, representation geometry, and robust AI systems, with experience in statistical model validation, Responsible AI, and high-dimensional scientific data analysis.

My work follows a research trajectory from statistical machine learning and model validation toward neural representation analysis and mechanistic interpretability.

Statistics
-> Statistical Machine Learning
-> Robust ML and Multi-Objective Model Validation
-> LLM Evaluation
-> Representation Geometry and Mechanistic Interpretability

Current Research

Beyond Shared Geometry: Discovering Approximate Group Representations in Neural Network Activations

My current research studies whether semantic transformations can be represented by approximately compositional operators acting on LLM hidden states, and whether this structure generalizes across prompts, concepts, and layers beyond simple geometric baselines.

The project focuses on transformer hidden representations, activation-space geometry, equivariance, operator learning, composition, inverse consistency, and statistical validation of representation structure. The goal is to determine whether transformations of meaning correspond to structured geometric transformations inside neural networks, while testing the robustness of the observed structure against geometric baselines and alternative explanations.

Ongoing research.

Robust Multi-Objective Challenger Search

The project combines XGBoost with mixed-variable NSGA-III to jointly search raw feature subsets and hyperparameters across predictive performance, directional residual disparity, and model sparsity. Experiments cover five public tabular benchmarks with multi-seed evaluation, threshold-sensitivity analysis, calibration diagnostics, and Pareto-frontier analysis.

Code

Research code release.

Representation Learning and Scientific ML

My earlier research involved statistical learning for high-dimensional scientific data, including co-fractionation mass spectrometry, protein-complex analysis, multiomics, clustering, automated data-analysis pipelines, and representation learning.

This work provides the foundation for my current interest in neural representations: how learned features organize information, how that organization can be measured, and how statistical validation can separate robust signal from artifacts.

Selected Publications

I have also contributed to published or accepted work involving cotton fiber genomics, transcriptomics, proteomics, multiomics, and high-dimensional biological data.

Some historical repositories on this account are preserved as stable research artifacts associated with published academic papers. Their structure is intentionally maintained for reproducibility and citation stability.

Background

  • Ph.D. in Statistics, Purdue University, 2025
  • Professional experience in statistical model validation and Responsible AI
  • Quantitative ML experience with model evaluation, robustness, and high-dimensional data
  • Research experience in statistical machine learning, scientific computing, and reproducible data-analysis pipelines

Pinned Loading

  1. Credit_Risk_Prediction_XGBoost Credit_Risk_Prediction_XGBoost Public

    Kaggle: Give Me Some Credit

    Jupyter Notebook

  2. R-code-S4_Class-protein-clustering-based-on-data-integration-of-corum-and-inparanoid R-code-S4_Class-protein-clustering-based-on-data-integration-of-corum-and-inparanoid Public

    R