The interesting claim is not “AI wrote hardware”
OpenTPU arrived in late September as an unusually inspectable FPGA accelerator project: an instruction-set compiler, a bit-exact Python simulator, SystemVerilog RTL, a host runtime, and working model demos live in one public repository. The author describes it as AI-developed and, just as importantly, as a learning project. That framing makes it useful. We can examine what evidence connects generated code to an executing board instead of debating an opaque demo.
The focused 30-day scan found a real developer cluster rather than broad consumer hype. The project reached 232 GitHub stars and eight forks at our October 7 check, while its main Hacker News discussion had roughly 240 points and 300 comments. A focused r/FPGA post had 131 votes and 29 comments. These are time-stamped attention signals, not quality certificates. The stronger signal is repository activity: a same-day commit repaired floating-point boundary behavior and documented model, unit, and timing results.
OpenTPU currently targets an Inspur YPCB-00338 card with a Xilinx Kintex-7 XC7K480T FPGA, two DDR3 channels, and a PCIe host connection. Its README reports execution of ten modern language models, including an int8 LFM2.5-230M path and a 4-bit Qwen3.5-0.8B path. The design is deliberately explicit: DMA, matrix, vector, and quantizer units are driven by instructions; there is no cache hiding data movement. That makes the performance bottleneck and the evidence boundary visible.
The useful question is not whether AI can produce RTL. It is whether the repository can prove that the compiled model, simulator, RTL, bitstream, and board executed the same intended computation.
That distinction matters beyond this project. “It generates the same sentence” is an appealing demo, but a language model can hide rare arithmetic errors. “Simulation passed” may mean a Python model passed, not synthesized RTL. “Runs at 133 MHz” may describe a timing target, not a closed design on the published commit. A trustworthy release binds each claim to a named artifact, method, and result.
Map the accelerator as one executable chain
The software starts from supported model weights and lowers operations into a compact instruction stream. Each instruction is represented by eight 32-bit words. A Python ISA simulator provides the golden behavioral path. The hardware implements the corresponding operations in SystemVerilog; the host runtime transfers programs and tensors over PCIe, coordinates external memory, and retrieves results.
model + tokenizer + quantization policy
-> kernel/compiler lowering
-> instruction image (8 x 32-bit words per instruction)
-> bit-exact Python ISA simulator
-> SystemVerilog decode and execution units
-> synthesized bitstream + timing constraints
-> PCIe host runtime + two-channel DDR3
-> logits -> host argmax -> next token
-> model-level output and hardware counters
The major units create natural test boundaries. DMA moves tensors between host or DRAM and local storage. The matrix unit performs the throughput-heavy dot products. Vector functions apply elementwise operations. The quantizer converts representations. Decode and control logic schedule those units. A defect in any one layer can look like a model-quality problem at the top, so the diagnostic record must preserve the lowest layer at which divergence first appears.
The repository reports a 133.33 MHz “Build B” with 17.1 GB/s theoretical DRAM bandwidth. Its measured int8 LFM2.5-230M decode used 14.5 GB/s—about 85% of that peak—and reached 59.0 device tokens per second, or 52.3 tokens per second wall-clock. Qwen3.5-0.8B at 4-bit was reported at 24.5 device and 23.3 wall-clock tokens per second. These numbers are meaningful only with the published method: greedy decode for 64 tokens after a 512-token prompt, host argmax in the loop, with device counters separated from wall time.
This architecture is pedagogically clean because it exposes data movement. A cacheless instruction machine makes transfers explicit and easier to account for, though not necessarily faster. If decode performance tracks sustained DRAM bandwidth, the next optimization is not automatically another multiply array. It may be weight packing, burst efficiency, double buffering, instruction overhead, or reducing round trips to host argmax.
Bind every claim to a release manifest
A reproducible accelerator release needs more than a commit and a screenshot. Model repositories change; compilers choose different kernels; synthesis tools alter placement; host software changes measurement boundaries. Record the inputs and outputs that let another engineer reconstruct the claim.
schema: accelerator-evidence/v1
projectCommit: "982d66817e..."
model:
repository: "vendor/model-name"
revision: "immutable-sha"
weightsSha256: "..."
tokenizerSha256: "..."
compile:
quantization: int8
compilerOptions: ["--target=build-b"]
instructionImageSha256: "..."
verification:
isaGolden: pass
directedNumericVectors: 5/5
rtlUnits: 211/211
differentialPrograms: pass
hardware:
board: "YPCB-00338"
part: "xc7k480t"
bitstreamSha256: "..."
clockMHz: 133.33
worstNegativeSlackNs: 0.012
benchmark:
promptTokens: 512
generatedTokens: 64
decoding: greedy
deviceTokensPerSecond: 59.0
wallTokensPerSecond: 52.3
decision: accepted_for_learning_benchmark
The decision label matters. Passing this contract would support “the pinned design reproduced the stated learning benchmark on the named board.” It would not support production reliability, safety certification, compatibility with arbitrary models, or a lifetime estimate. Scope the conclusion to the tests actually run.
| Claim | Required evidence | Common false substitute |
| Compiler is correct | Known kernels plus randomized differential comparison against a reference | One model produces plausible text |
| RTL matches the ISA | Same instruction and initial state produce matching intermediate and final state | Separate unit tests that never meet |
| Bitstream meets timing | Post-route report for the released constraints and bitstream | Requested clock in a build script |
| Board output is correct | Captured board tensors or logits compared under declared tolerance | Same final token sequence |
| Performance is reproducible | Warm-up, workload, counters, repetitions, clock, model hashes, and wall boundary | A peak number without method |
| Project is production-ready | Reliability, security, recovery, compatibility, and operational qualification | Stars, demos, or a passing test suite |
Use a verification ladder, not one end-to-end pass
Start with pure functions: instruction encoding, address calculations, quantization, rounding, saturation, special values, and each arithmetic primitive. Then test units with cycle-aware interfaces. Next, run short instruction programs through both the ISA simulator and RTL, comparing architectural state after every instruction or synchronization point. Only after those layers agree should a full model become an acceptance test.
def verify(program, initial_state, tolerance):
golden = isa_simulator.run(program, initial_state, trace=True)
rtl = verilated_rtl.run(program, initial_state, trace=True)
for step in synchronization_points(program):
compare(rtl.memory(step), golden.memory(step), exact=True)
compare(rtl.integer_state(step), golden.integer_state(step), exact=True)
compare(rtl.float_state(step), golden.float_state(step),
atol=tolerance.atol, rtol=tolerance.rtol,
nan_policy="same_class")
return first_divergence(rtl, golden)
Compare integers and instruction-visible memory exactly. Floating-point comparison needs an explicit policy: data type, rounding mode, handling of signed zero, infinity, NaN class or payload, absolute and relative tolerance, and whether accumulation order may differ. Do not pick a tolerance after seeing a failure. Version it with the operation and representation.
Random testing is valuable when it is constrained by real preconditions. Generate legal shapes, aligned and misaligned addresses when supported, minimum and maximum vector lengths, overlapping buffers, dependency chains, and back-pressure patterns. Add every discovered defect as a permanent small regression. Model tests then verify integration: tokenizer, embeddings, compiler lowering, memory plan, kernel sequence, board transfer, logits, and decode loop.
ChipVerilog, a recent research system for Verilog generation, illustrates the broader pattern: generation is paired with compilation and simulation rather than accepted from text alone. The exact tool can vary—Verilator is a practical open-source route for compiling SystemVerilog models—but the contract should not. Generated RTL enters the same review, lint, simulation, equivalence, synthesis, and timing pipeline as human-written RTL.
A same-day fix shows why token tests are insufficient
The most instructive OpenTPU evidence on October 7 was not a benchmark. Commit 982d66817e fixed FP32 boundary, NaN, and argmax behavior. The commit record says the previous production path failed four of five directed cases, while the revised RTL and board passed all five. It also records 44 fast qualification passes, 211 RTL unit passes, 1,056 non-RTL test passes, and positive timing slack.
This is not an embarrassment; it is the verification story. A large general suite and plausible model output coexisted with arithmetic corner-case defects. The directed vectors exposed behavior at representation boundaries where ordinary activations may rarely land. Argmax makes those details consequential: mishandling NaN or a comparison tie can select a different token even when most logits are close enough.
Build a numerical suite from partitions, not anecdotes. Include zero and negative zero; smallest subnormal and normal values; maximum finite values; positive and negative infinity; quiet and signaling NaN if represented; equal values; adjacent representable values; values on quantization thresholds; long accumulations; overflow and underflow; and permutations that test whether the result depends improperly on operand order.
Reference semanticsName the ISA simulator or mathematical function that defines expected behavior.
Directed boundariesExercise encoding limits, special values, ties, rounding transitions, and saturation.
Random differentialGenerate legal programs and stop at the first state divergence.
Model witnessCompare selected intermediate tensors and final logits, not tokens alone.
Board replayRun the exact vectors on the released bitstream and preserve raw results.
Separate device throughput from the experience a user sees
OpenTPU usefully reports device decode and wall-clock decode separately. The difference includes host argmax, transfers, synchronization, and runtime overhead. Both numbers answer legitimate questions. Device throughput helps locate hardware bottlenecks; wall throughput describes the current application path. Publishing only the larger number would obscure engineering work still outside the FPGA.
Prefill and decode also stress different paths. Prefill processes many prompt tokens in parallel and often uses compute efficiently. Autoregressive decode repeatedly streams weights for one next token and can become memory-bandwidth bound. Report prompt length, generated length, batch size, quantization, clock, decoding policy, model revision, warm-up, repetitions, percentile or spread, and whether tokenization is inside the timer.
For bandwidth claims, show the denominator. The repository's 17.1 GB/s peak follows the chosen memory interface and clock; the measured 14.5 GB/s is a counter-derived workload result. “85% utilization” is informative because both are stated. Also report useful bytes separately from protocol or reread traffic when counters allow it. A design can saturate memory and still waste bandwidth.
Never compare a learning FPGA directly with a commercial accelerator using tokens per second alone. Model architecture, parameter count, precision, quality, prompt length, batch, compiler maturity, power, memory capacity, and price differ. The fair comparison is internal: did a specific change preserve numerical acceptance while improving a pinned workload and maintaining timing?
Reproduce the software path before touching a board
Use a dedicated environment, inspect the repository, and pin an immutable commit. The project currently documents an editable Python install, optional test dependencies, a Hugging Face model download, an ISA-backed chat path, and a separate board setup. Commands change, so treat these as a starting map and verify the current README.
git clone https://github.com/FeSens/openTPU.git
cd openTPU
git checkout <reviewed-commit-sha>
python3 -m venv .venv
source .venv/bin/activate
pip install -e .
pip install pytest torch transformers
python3 -m pytest -q
# The RTL suite requires Verilator 5 according to the repository.
# After downloading a supported pinned model revision:
otpu-chat --model lfm2 --backend isa
Do not start with make bit. First reproduce the reference simulator, run non-RTL and RTL tests, inspect compiler output for a tiny kernel, and save a known trace. For board work, verify the exact FPGA part and card revision, toolchain version, constraints, PCIe permissions, reset procedure, and recovery path. Then synthesize, review utilization and timing, program through the documented JTAG path, run the setup utility with the minimum required privilege, and replay directed vectors before a language model.
Keep model downloads and build inputs immutable. A repository name is not a revision. Record file hashes, licenses, and any conversion or quantization command. The OpenTPU code is Apache-2.0, but model weights and FPGA toolchains have separate terms. Do not infer that the repository license covers every input.
Failure modes that a polished demo can hide
| Failure | Why it slips through | Control |
| Same tokens, wrong logits | Argmax stays unchanged on the demo prompt | Compare intermediate tensors and logits under a declared policy |
| Simulator and RTL share the same bug | Both derive from one mistaken interpretation | Independent spec vectors and a third reference for critical operations |
| Test passes, timing fails | Behavioral simulation has no routed delays | Bind the post-route report and constraints to the bitstream |
| Peak presented as measured | Theoretical bandwidth looks authoritative | Publish raw counters, workload, repetitions, and both values |
| Host work omitted | Device timer excludes transfers and argmax | Report device and wall-clock boundaries side by side |
| Moving model dependency | A model name resolves to new files | Pin revisions and hashes for model, tokenizer, and conversions |
| AI provenance becomes an exemption | Generated code is treated as experimental | Apply the normal HDL review and verification pipeline |
Release checklist
- Pin repository, submodules, model, tokenizer, compiler, simulator, FPGA tools, and board revision.
- Save the instruction image, bitstream, timing report, utilization report, logs, traces, and hashes.
- Run instruction encoding, numerical boundary, unit, constrained-random, ISA-to-RTL differential, and model tests.
- Compare board results with the golden path at a level deeper than generated tokens.
- Declare float and quantization tolerances before evaluating results.
- Report prefill and decode separately, with device and wall-clock timing.
- Preserve raw counters and explain peak-versus-measured bandwidth.
- Record every exception, owner, disposition, and test rerun after a fix.
- Use a scoped conclusion such as learning benchmark, not production readiness.
FAQ
What is OpenTPU?
It is an open-source learning accelerator stack with compiler tooling, a bit-exact Python ISA simulator, SystemVerilog RTL, a host runtime, and an FPGA implementation. It is unrelated to treating “TPU” as proof of compatibility with Google's production hardware.
Does matching generated text prove correctness?
No. It proves a useful end-to-end example. Many incorrect logits produce the same winning token, and a rare arithmetic boundary may never occur. Add directed numerical tests, differential traces, RTL units, synthesis evidence, and board replay.
Why use an ISA simulator?
It gives the compiler and RTL a shared architectural contract. It is faster to inspect than hardware, can emit exact traces, and supports first-divergence debugging. Critical semantics still need independent specification vectors so the simulator does not merely duplicate the RTL's mistake.
Is the reported performance comparable to a GPU?
Not from tokens per second alone. Model, precision, batching, prompt, power, memory, software, and timing boundaries differ. Use the published numbers to reproduce this build and compare controlled revisions of the same workload.
Sources and further reading
- OpenTPU repository and README — architecture, supported workflows, board target, tests, benchmarks, license, and current project scope.
- OpenTPU FP32 boundary fix — directed cases, qualification results, RTL and non-RTL test counts, and timing evidence for the October 7 repair.
- r/FPGA OpenTPU discussion — focused community discussion; used as attention and practitioner context, not correctness evidence.
- Hacker News OpenTPU discussion — 241 points and 303 comments at the October 7 check; discussion volume is not treated as validation.
- ChipVerilog — recent research on generated Verilog paired with compilation and simulation feedback.
- Verilator documentation — current reference for compiling and testing SystemVerilog models.
- AMD Kintex-7 documentation — vendor documentation for the FPGA family used by the target card.
- In-datacenter performance analysis of a TPU — foundational context for workload-oriented accelerator evaluation.
Accessed October 7, 2026. Repository counts and benchmarks are time-sensitive. Performance figures are attributed to the OpenTPU project and were not independently reproduced for this article. The focused scan used GitHub, Hacker News, Reddit, and web cross-checks; unrelated social matches were excluded.