The Chip Design Bottleneck
AI models evolve at a brutal pace. Software updates roll out weekly, and architectures pivot every few months. Yet the underlying silicon is locked in years before those models ever see production. You spend millions of dollars and countless engineering hours betting on what workloads will look like half a decade down the line. If your assumptions miss the mark, you pay for it twice: once for the general-purpose safety margins baked into the die, and again when real-world workloads map poorly onto frozen hardware.
That mismatch has crippled hardware development for decades. According to a functional verification study commissioned from Wilson Research Group by Siemens EDA, only 14 percent of integrated circuit and ASIC projects achieved first-silicon success in 2024, marking the lowest rate in two decades. Three-quarters of projects fell behind schedule.
Architect Labs thinks it has a way out. The startup’s AI platform recently designed, verified, programmed, and deployed a low-power inference chip called Redwood from scratch in under two weeks.
From Specification to FPGA in Fourteen Days
Rather than accelerating isolated tasks within traditional design workflows, the Architect Labs Platform attempts to reimagine the entire pipeline. In their published experiments, two human architects provided a high-level specification for Redwood. From there, the AI system took over.
The platform handled performance modeling, generated digital hardware descriptions, tested and verified the design, and produced the necessary firmware and kernels. Every block reached 95 percent code and functional coverage without human verification engineers lifting a finger. According to the researchers, the first RTL design sent from simulation to the field-programmable gate array (FPGA) contained zero bugs.
When the team needed to make specification changes, the turnaround was even faster. Once updates were verified, a modified design was back running on the FPGA in under 48 hours—a stark contrast to conventional flows where design stages are rigidly frozen and iterative changes require ad-hoc workarounds or get shelved for the next generation.
Microarchitectural Exploration at Scale
Building efficient hardware requires exploring a vast design space. Human engineering teams are naturally constrained by time and bandwidth, forcing them to settle on promising architectures early and abandon alternatives.
The Architect Labs system operates differently. It can explore a microarchitectural search space an order of magnitude larger than a human team could manage in the same window. For one vector engine, the system generated and verified multiple design options with varying control logic and data paths, grinding away over several days to optimize specifically for area and timing constraints.
Human experts retain control over the high-level specification, adjusting parameters based on functional, area, performance, timing, and power feedback. The system then regenerates and reverifies the design from that updated baseline. By collapsing the software-to-silicon stack into a single optimization loop, hardware and software are co-designed and verified under one unified objective.
Architecture of the Redwood Accelerator
Redwood was built to tackle three core hardware headaches: latency, power consumption, and execution predictability.
To curb latency, the chip keeps data as close to the compute units as possible. Instead of repeatedly pulling data from shared memory, Redwood relies on a grid of identical building blocks called tiles. Each tile executes a specific slice of work and passes the results directly to its neighbors in a planned data flow pattern.
Power efficiency is achieved by separating control from compute, allowing logic to slow down or power off when idle. Near-memory compute and intelligent flow control on the on-chip network minimize wasted energy. Furthermore, a global timer schedules each kernel precisely, bringing deterministic predictability to inference workloads.
Measured FPGA Results Versus Silicon Projections
Evaluating an unbuilt chip always demands a healthy dose of skepticism. Redwood currently runs exclusively on an FPGA; it has not yet been physically fabricated in silicon.
To benchmark performance, the researchers evaluated a configuration called Redwood Nano against Nvidia’s Jetson Orin Nano, serving as the commercial baseline for edge AI. Measured on the FPGA running at 250 megahertz, Redwood Nano achieved an average throughput of 12.1 tokens per second on the Qwen3 0.6B large language model. By comparison, the Jetson Orin Nano hit 28 tokens per second at a 1020 megahertz clock speed.
The projected advantages over Nvidia emerge when looking ahead to fabricated silicon. Using a Samsung 8nm-class process comparable to the Jetson's, the researchers project that Redwood Nano would reach 49 tokens per second—about 1.75 times the Jetson's measured throughput—while consuming roughly half the power (1.335 watts versus 2.59 watts). That translates to a 3.4-fold improvement in performance per watt, with an NPU block occupying a mere 2.88 square millimeters.
Recursive Self-Improvement in Hardware
In one of the more intriguing experiments, the team deployed the Qwen3 model directly onto Redwood, exposed it as an inference endpoint within their AI system, and sampled it repeatedly. The model subsequently identified multiple timing improvements and kernel optimizations for several of its own operations.
The researchers describe this as an early, narrow instance of recursive self-improvement: an AI system designing an accelerator, deploying a model on it, and using that model to optimize subsequent generations of the hardware.
Architect Labs is now working toward a physical tapeout of Redwood at TSMC. If the design successfully translates from FPGA emulation to fabricated silicon, the most significant takeaway might not be Redwood itself, but the newfound ability for hardware teams to iterate alongside fast-moving AI models rather than guessing the future years in advance.