Curriculum Vitae
A PDF version of the CV is available here. Contact: peli@stanford.edu.
Education
Ph.D. in Electrical Engineering, Stanford University
- Entering the third year in Fall 2026
- Research focus: the boundary between compilers and hardware — compile-time reasoning about runtime behavior, data lifetime profiling, and solver-backed memory placement for domain-specific accelerators
M.S. in Electrical Engineering, Stanford University, expected December 2026
- Matriculated concurrently with the Ph.D. program; remaining degree requirements are being completed in Fall 2026
B.S.E. in Computer Science, University of Michigan, May 2024
- Graduated summa cum laude
- Minor in Civil Engineering
- Enrolled in the College of Engineering Honors Program
- GPA: 4.0/4.0 with 144 completed credits
Honors & Awards
- Qualcomm Innovation Fellowship (North America) — Finalist (April 2026), for the proposal “Refresh-Aware Scheduling and Memory Planning for AI Accelerators with Short-Term RAM” (project Remi), with collaborators Nathan Kaplan and Prof. Sharad Malik, Princeton University.
- Sunlin and Priscilla Chou Graduate Fellowship (September 2024), first-year Ph.D. fellowship at Stanford University.
Work experience
- AI Hardware Research Intern, TSMC Corporate Research, San Jose, CA (June – September 2026)
- Extended an in-house analytical model for ASIC and GPU performance to support workload profiling for large language model inference, and investigated embedded DRAM opportunities across data types including activations, KV cache, and weights.
- Built an analytical cost model for multi-chip scaling behavior, unifying several interconnect topology families under a single latency and energy notation.
- This work is proprietary; its detailed results and architectural specifics are not disclosed here.
- Head Teaching Assistant, CS 217: Hardware Accelerators for Machine Learning,
Stanford University, Stanford, CA (Winter 2026)
- Rebuilt the course from the ground up with three other teaching assistants and two faculty instructors after a four-year hiatus; enrollment reached roughly 100 students against an expected 40 to 60.
- Authored the lab sequence in which students build a large language model accelerator incrementally — systolic array, memory hierarchy, and compiler backend — culminating in an open-ended final project. All starter code was written from scratch; the SystemC and Verilog accelerator code was adapted from my own 16nm tapeout with block floating-point and microscaling formats removed.
- Tuned DSP utilization for the arithmetic circuits through Siemens Catapult and Xilinx Vivado to make FPGA lab turnaround tractable, which was the binding constraint on the lab design.
- Designed the final-project rubric and gave in-person feedback on every project proposal.
- Faculty instructors: Prof. Thierry Tambe and Prof. Kunle Olukotun.
- Research Mentor, Stanford Engineering Undergraduate Visiting Research Program,
Stanford, CA (Summer 2025)
- Defined the problem scope for a rising senior's ten-week summer research project and onboarded them to two active projects.
- Their deliverable extended the Synopsys ASIP Designer commercial toolchain with cycle-accurate memory-access logging, which the toolchain lacked, by instrumenting the RTL testbench memory module to emit load and store records with cycle numbers and addresses.
- They also built a Python pipeline converting raw traces into a GainSight-compatible memory log, and designed a DMA-based DRAM subsystem for a VLIW and RISC-V AI core including a three-instruction custom ISA extension and an HBM-parameterized latency model.
- Four processor designs were profiled across workloads from single CNN layers up to Llama-3-8B, plus quicksort, motion estimation, and FFT. Presented at the program's closing symposium, August 2025.
- Research Assistant, Stanford University, Stanford, CA (September 2024 – March 2026)
- Heterogeneous on-chip memory systems for AI accelerators, within the Differentiated Access Memories project. Built GainSight, a data lifetime profiling and memory composition framework.
- On the GEMMA 16nm AI accelerator tapeout (targeting Mamba-class state-space model inference), joined roughly four months before tapeout and redesigned the embedded DRAM partition, the design-for-test module, and the refresh controller, whose refresh policy is driven by data-lifetime patterns identified through GainSight profiling. Also implemented the MXFP6 quantization and arithmetic logic, authored the AXI and DMA transaction specifications, and designed a high-throughput SIMD compute unit using high-level synthesis.
- Software Test Engineer Intern, ASML Silicon Valley, San Jose, CA (May – August 2023)
- Owned the automation of nightly regression triage for focus-exposure modeling software. The suite ran roughly 300 test cases over about six hours.
- Attributing each failing test to a commit and an owner previously cost two to three hours per day across a five-engineer team; afterwards attribution ran unattended in five to ten minutes and emailed a summary of failing tests with their owners, leaving only the triage decision to a human.
- The scripts remain an actively maintained part of the team's workflow three years later.
- Supervisor: Mingjing Zhao.
- Instructional Aide, CEE 375: Sensors, Circuits, and Signals, Department of Civil and
Environmental Engineering, University of Michigan, Ann Arbor, MI (January – April 2023)
- Ran weekly laboratory sections in which students built sensing and signal-processing circuits on breadboards using Arduino and MATLAB, supporting them on both the mathematics and the physical debugging.
- Faculty instructor: Dr. Jeff Scruggs.
- Research Assistant, University of Michigan Transportation Research Institute,
Ann Arbor, MI (February – August 2022)
- Co-developed a computer vision system for monitoring occupant body posture in Level-3 autonomous vehicles, and built multi-body occupant models in MATLAB and Mathematica from 3D body-scan data for the HumanShape project.
- Supervisors: Dr. Monica Jones, Dr. Jingwen Hu.
- Software Engineering Intern, Dell EMC, Shanghai, China (June – August 2021)
- Automated lifecycle-management validation for hyperconverged infrastructure across the Kubernetes interface, the VMware ESXi hypervisor, and VMware Cloud Foundation, using Python and Selenium.
- A manual procedure that took three to four hours of stepping through routines by hand became roughly 30 minutes running unattended, and eliminated most physical trips to the server room to check cluster status.
- Also interfaced with C++ kernel code in the bare-metal hypervisor and debugged physical server racks when experimental deployments failed to boot.
- Supervisor: Carl Shi.
Skills
- Software: Python (PyTorch, JAX), C/C++, CUDA, LLVM/MLIR, Z3 SMT solver, RISC-V/ARM/x86 assembly
- Hardware: SystemVerilog, Verilog, SystemC/TLM, AXI, high-level synthesis (Siemens Catapult, Xilinx Vivado), Synopsys ASIP Designer, logic synthesis, design for test
- Tools: Git, Docker, Linux administration, Slurm, CI/CD, Jenkins, Atlassian tools, qTest, Selenium
- Spoken languages: English, Mandarin Chinese
Projects
- GainSight — data lifetime profiling for heterogeneous on-chip memory — Research project, first author and lead developer
- Flan — solver-backed memory placement in MLIR — Research project, originator and lead
- GEMMA — 16nm AI accelerator tapeout — Silicon project, memory and design-for-test blocks
Publications
- LLM-FSM: Scaling Large Language Models for Finite-State Reasoning in RTL Code Generation — Yuheng Wu, Berk Gokmen, Zhouhua Xie, Peijing Li, Caroline Trippel, Priyanka Raina, and Thierry Tambe. 2026. LLM-FSM: Scaling Large Language Models for Finite-State Reasoning in RTL Code Generation. https://doi.org/10.48550/arXiv.2602.07032
- Towards Memory Specialization: A Case for Long-Term and Short-Term RAM — Peijing Li, Muhammad Shahir Abdurrahman, Rachel Cleaveland, Sergey Legtchenko, Philip Levis, Ioan Stefanovici, Thierry Tambe, David Tennenhouse, Caroline Trippel, and H.-S. Philip Wong. 2025. Towards Memory Specialization: A Case for Long-Term and Short-Term RAM. In Workshop on Disruptive Memory Systems (DIMES '25), October 13, 2025. Association for Computing Machinery, Seoul, Korea (South), 10. https://doi.org/10.1145/3764862.3768175
- The Future of Memory: Limits and Opportunities — Samuel Dayo, Shuhan Liu, Peijing Li, Philip Levis, Subhasish Mitra, Thierry Tambe, David Tennenhouse, and H.-S. Philip Wong. 2025. The Future of Memory: Limits and Opportunities. https://doi.org/10.48550/arXiv.2508.20425
- GainSight: A Unified Framework for Data Lifetime Profiling and Heterogeneous Memory Composition — Peijing Li, Matthew Hung, Yiming Tan, Konstantin Hoßfeld, Jake Cheng Jiajun, Shuhan Liu, Lixian Yan, Xinxin Wang, Philip Levis, H.-S. Philip Wong, and Thierry Tambe. 2025. GainSight: A Unified Framework for Data Lifetime Profiling and Heterogeneous Memory Composition. https://doi.org/10.48550/arXiv.2504.14866
- OpenGCRAM: An Open-Source Gain Cell Compiler Enabling Design-Space Exploration for AI Workloads — Xinxin Wang, Lixian Yan, Shuhan Liu, Luke Upton, Zhuoqi Cai, Yiming Tan, Shengman Li, Koustav Jana, Peijing Li, Jesse Cirimelli-Low, Thierry Tambe, Matthew Guthaus, and H.-S. Philip Wong. 2025. OpenGCRAM: An Open-Source Gain Cell Compiler Enabling Design-Space Exploration for AI Workloads. https://doi.org/10.48550/arXiv.2507.10849
- A Communication Protocol for Securing Connected Vehicle Platoons Using Joint Hardware-Software Means — Peijing Li and Neda Masoud. 2023. A communication protocol for securing connected vehicle platoons using joint hardware-software means. In 6th Student Poster Competition at the CCAT Global Symposium, April 05, 2023. Center for Connected and Automated Transportation, Ann Arbor, MI. Retrieved from https://ccat.umtri.umich.edu/symposium/2023-symposium/#poster
Academic Projects
- "Patchouli": Performance Analysis and Tiling Choice Optimization Using LLVM IR — Course project, EECS 583 Advanced Compilers, University of Michigan
- A tour of accelerator architectures for AI/ML applications — Topic lecture, EECS 573 Microarchitecture, University of Michigan
- An out-of-order, super-scalar implementation of the RISC-V ISA in the style of the P6 micro-architecture — Course project, EECS 470 Computer Architecture, University of Michigan
- Investigation of representation of accidents in car-following models — Course project, CEE 551 Traffic Science, University of Michigan
- Implementation and evaluation of a nested, variable-depth UNet++ model architecture for medical imaging segmentation — Course project, EECS 442 Computer Vision, University of Michigan
Service & Leadership
- Props & Uniforms Staff, Leland Stanford Junior University Marching Band, Stanford, CA (March 2025 – Present)
- Undergraduate Student Member, Tau Beta Pi Engineering Honor Society, Michigan Gamma Chapter, Ann Arbor, MI (September 2022 – May 2024)
- Student Life Committee Member, University of Michigan Engineering Student Government, Ann Arbor, MI (September 2020 – May 2022)