LLM-FSM: Scaling Large Language Models for Finite-State Reasoning in RTL Code Generation
Recommended citation: Yuheng Wu, Berk Gokmen, Zhouhua Xie, Peijing Li, Caroline Trippel, Priyanka Raina, and Thierry Tambe. 2026. LLM-FSM: Scaling Large Language Models for Finite-State Reasoning in RTL Code Generation. https://doi.org/10.48550/arXiv.2602.07032
View paper
Abstract
Finite-state reasoning, the ability to understand and implement state-dependent behavior, is central to hardware design. In this paper, we present LLM-FSM, a benchmark that evaluates how well large language models (LLMs) can recover finite-state machine (FSM) behavior from natural-language specifications and translate it into correct register transfer-level (RTL) implementations. Unlike prior specification-to-RTL benchmarks that rely on manually constructed examples, LLM-FSM is built through a fully automated pipeline. LLM-FSM first constructs FSM with configurable state counts and constrained transition structures. It then prompts LLMs to express each FSM in a structured YAML format with an application context, and to further convert that YAML into a natural-language (NL) specification. From the same YAML, our pipeline synthesizes the reference RTL and testbench in a correct-by-construction manner. All 1,000 problems are verified using LLM-based and SAT-solver-based checks, with human review on a subset. Our experiments show that even the strongest LLMs exhibit sharply declining accuracy as FSM complexity increases. We further demonstrate that training-time scaling via supervised fine-tuning (SFT) generalizes effectively to out-of-distribution (OOD) tasks, while increasing test-time compute improves reasoning reliability. Finally, LLM-FSM remains extensible by allowing its FSM complexity to scale with future model capabilities.
Citation
The BibTeX entry for this paper is
@misc{wu_llm-fsm_2026,
title = {{LLM}-{FSM}: {Scaling} {Large} {Language} {Models} for {Finite}-{State} {Reasoning} in {RTL} {Code} {Generation}},
shorttitle = {{LLM}-{FSM}},
url = {http://arxiv.org/abs/2602.07032},
doi = {10.48550/arXiv.2602.07032},
publisher = {arXiv},
author = {Wu, Yuheng and Gokmen, Berk and Xie, Zhouhua and Li, Peijing and Trippel, Caroline and Raina, Priyanka and Tambe, Thierry},
month = feb,
year = {2026},
note = {arXiv:2602.07032 [cs.AI]},
keywords = {Computer Science - Artificial Intelligence, Computer Science - Hardware Architecture, Computer Science - Computation and Language},
}