openTPU brings AI inference to a Kintex-7 FPGA
FeSens’s open-source project combines an FPGA accelerator with the simulator, compiler and host software used to run and profile it.
FeSens’s openTPU project packages an AI accelerator and its supporting tools in one open-source repository. The hardware targets an Inspur PCIe card built around a Xilinx Kintex-7 FPGA, and the project says it runs modern AI models on the card using their real weights.
The repository includes the SystemVerilog hardware design, an instruction set, a bit-exact simulator, a kernel language and compiler, and host software. A profiler can record runs from the RTL, simulator or card and show activity by instruction and hardware unit.
FeSens says the card has run ten models and produced the same tokens as the simulator, bit for bit. The project also reports that its newer production build decodes several models 8–9% faster than an earlier build, with DRAM bandwidth reaching 91–94% of peak, compared with 82–87%.
For Qwen3.5 and other models, the project says 4-bit weights raise decoding speed by 40–45%, while reducing data per token by about a third. It also reports a measurable cost in perplexity. Models larger than the card’s 4 GiB memory can run with experts streamed from host storage.
The performance and compatibility figures are FeSens’s own measurements and have not been independently verified in the supplied material. The project describes openTPU as developed by AI, but its description does not detail the AI agents’ role or the extent of human involvement.
