Skip to content
AI.info

Research

LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons

Overview Research area: Efficient machine learning / hardware-software co-design for vision transformers (ViTs), specifically FPGA-based edge inference. Technical level: Advanced — the paper assumes f

LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons
arXiv
2511.00812
Published
2025-11-02
Authors
Shashank Nag, Alan T. L. Bacellar, Zachary Susskind, Anshul Jha, Logan Liberty, Aishwarya Sivakumar, Eugene B. John, Krishnan Kailas, Priscila M. V. Lima, Neeraja J. Yadwadkar, Felipe M. G. Franca, Lizy K. John

AI summary

Overview

Research area: Efficient machine learning / hardware-software co-design for vision transformers (ViTs), specifically FPGA-based edge inference.

Technical level: Advanced — the paper assumes familiarity with transformer architectures, quantization, FPGA resource primitives (LUTs, DSPs, BRAMs), and look-up-table neuron networks.

One-sentence scope: The paper replaces the multiplication-heavy MLP (channel mixer) blocks inside a quantized vision transformer with learned look-up-table neurons and builds a matching FPGA accelerator, reporting comparable accuracy with smaller models and better energy efficiency.

What This Paper Is About

Vision transformers deliver strong accuracy but are costly to run on edge hardware: even a compact model like DeiT-T needs roughly 20 MB of weight storage (about 9100 BRAM18s), more than many FPGAs have on chip. The authors observe that MLP/channel-mixer layers account for over 60% of model weights and 55% of MAC operations (a range of 50–70% of MACs across ViTs), so those layers are the bottleneck. Their goal is a transformer whose channel mixing is done by learned look-up-table neurons instead of multiplications, paired with an FPGA accelerator that keeps all weights on chip.

Key Contributions

  1. A novel learnable LUT-based channel-mixer block, including a differentiable "conditional summation" output layer, that integrates into a standard transformer encoder in place of the MLP.
  2. An edge-oriented vision transformer design, LL-ViT, that pairs the LUT channel mixer with a conventional multi-head self-attention token mixer, keeping the model fully differentiable and trainable end-to-end.
  3. A co-designed low-power FPGA accelerator that avoids moving weights off chip, with a dedicated processing element for the LUT channel mixer, evaluated against prior ViT accelerators.
  4. The first application of LUT/logic-based neural layers to a vision transformer backbone and to relatively complex vision tasks (CIFAR-10, CIFAR-100, Tiny-ImageNet, Flowers-102), where prior LUT models had reported much lower CIFAR-10 accuracy.

Main Findings

  • Accuracy matches the quantized baseline: LL-ViT reaches 95.5% on CIFAR-10 (baseline I-ViT-T: 95.4%), 78.8% on CIFAR-100 (baseline 79.2%), 60.9% on Tiny-ImageNet (baseline 60.4%), and 91.6% on Flowers-102 (baseline 91.3%) per Table II; the results text cites 91.3% for Flowers-102. The authors used the default train-test splits of these datasets, resized images to 224x224, and applied the DeiT data augmentations.

  • Smaller models and fewer multiplies: The paper reports eliminating over 60% of model weights and 50% of the multiplications. Table II lists model size dropping from 5.06 MB to 1.93 MB; Table III lists parameters dropping from 5.3 M to 2.5 M and GMACs per inference from 1.31 to 0.65.

  • Energy and latency gains: LL-ViT achieves 1.9x better energy efficiency (4.05 mJ to 2.14 mJ per inference) and 1.3x lower latency (6.93 ms to 5.33 ms) than the integer-quantized I-ViT-T baseline accelerator on the same FPGA. The authors call the energy estimate conservative because it excludes off-chip fetching of baseline MLP weights.

  • On-device throughput and power: On an AMD Xilinx xcvu9p-flgb2104-2-i (Virtex series) FPGA at a 200 MHz target, LL-ViT runs at 1083 FPS within a 10.9 W power budget, using 0 DSPs, 589K LUTs, 229K flip-flops and 1425 BRAM36s, for 99.35 FPS/W.

  • Better than prior LUT-based models on CIFAR-10: Prior LUT-neuron and logic-based works report CIFAR-10 accuracy in the 40–80% range: TreeLUT 42.9%, DiffLogicNet 57.3%, DWN 57.5%, FINN 80.1%, none of which implement a ViT architecture. LL-ViT reports 95.5% with a 1976.3 KiB parameter size and 587K LUTs.

  • Layerwise energy shift: For one encoder layer per sample, total energy falls from 338.10 uJ to 178.2 uJ; the two baseline MLP dense layers each cost 94.35 uJ, while the LL-ViT channel mixer costs 28.8 uJ.

  • Post-training quantization of encoded values: Accuracy is 95.6% at int8, 95.5% at int4 and 90.3% at int2; int4 was chosen as the best tradeoff.

  • Robustness to baseline quantization and scaling: LL-ViT shows over 1.8x energy efficiency regardless of whether the baseline uses 4-bit, 2-bit or binary quantization, and the energy savings hold as latent dimension and image size vary.

  • Optimizing an already compact model: Applied to the CCT-2/3x2 backbone from Compact Convolutional Transformers, LL-ViT reaches 87.4% on CIFAR-10 while cutting MAC operations by about 1.5x.

  • Comparison to other FPGA ViT accelerators: At comparable resource usage (16x16 systolic arrays for this comparison), LL-ViT reports higher FPS/W than most prior works; the authors note HG-PIPE is the exception in FPS but consumes 4x the power, and that HG-PIPE uses 3-bit quantization while LL-ViT uses 8-bit.

Methodology in Plain English

The authors first profile vision transformers to find where the cost lives, and confirm the feed-forward MLP blocks dominate both weights and multiply-accumulate operations. Instead of accelerating those MLPs, they swap them out. Each channel-mixer MLP is replaced by a block that begins with a thermometer encoding layer (which turns activations into bit sequences), continues through one or more layers of LUT neurons, and ends with a conditional summation layer.

A LUT neuron concatenates its inputs and "looks up" its output rather than multiplying and adding. Making these neurons trainable is the hard part: the authors adopt the Extended Finite Difference gradient technique from Differentiable Weightless Neural Networks so gradients can flow through the table entries and inputs. Because a channel mixer sits in the middle of a transformer, it must emit real-valued features across all channels, so the conditional summation layer adds a learned, full-precision encoded value whenever a final-layer LUT output is 1 and skips it when the output is 0 — no multiplications, and smooth gradients during training. Those encoded values are quantized after training to match the rest of the network.

The optimal configuration found was two LUT layers with (768, 192) LUTs and 8-bit thermometer encoding, with the encoded values quantized to 4 bits. Models were built in custom PyTorch classes, trained on an A100 GPU using the baseline's learning rate, scheduler and optimizer, with attention layers left as standard multi-head self-attention because token mixing is not weight-dominated.

For hardware, the authors generate SystemVerilog RTL (with the processing element RTL generated from the trained model using custom mako scripts) and evaluate a design where every layer gets its own dedicated block and frames are pipelined through the encoder stack. A 32x32 systolic array handles the attention matrix multiplications and interfaces with the channel-mixer PE, which processes all channels of a row in parallel using ping-pong buffers to avoid stalls. Thermometer encodings are built from comparator blocks, LUT neurons map onto the FPGA's logic slices, and the conditional summation uses a 2:1 MUX plus an adder per latent dimension. All weights are read into dedicated BRAMs at initialization and stay on chip. Synthesis targets the xcvu9p-flgb2104-2-i device at 200 MHz in Vivado; power and energy are estimated with a 12.5% switching activity factor using AMD Power Design Manager.

Why This Matters

This work argues that learned LUT neurons are not just an inference-time trick for speeding up multiplications — they can serve as native building blocks inside a larger, complex model such as a vision transformer, and can outperform prior LUT/logic network approaches on realistic vision benchmarks.

Real-world applications:

  • Battery-operated robotics and autonomous systems that need low-latency vision inference under tight power budgets.
  • Smart cameras and industrial inspection devices where on-chip-only weights eliminate off-chip memory traffic and its energy cost.
  • Latency-sensitive edge perception in drones or embedded vision modules where 10.9 W-class power and high frame rates matter.
  • FPGAs deployed in settings where model architectures change quickly and reconfigurable hardware is preferable to fixed-function ASICs.

Industry relevance: The results speak directly to the FPGA and edge-AI accelerator market, where prior ViT accelerators either stream weights from off-chip memory or run at high power (the paper cites hybrid Versal approaches in the tens of watts). An accelerator using zero DSPs and fitting entirely on chip, at 1083 FPS and 10.9 W, is a concrete alternative for vendors targeting power-constrained deployments.

Future Directions

  • Extending LL-ViT beyond the tiny ViT baseline to larger vision transformer variants, which the authors say would show similar hardware improvements given the accelerator's structural composition.
  • Combining the LUT channel mixer with aggressive quantization methods such as BinaryViT, since the ablation suggests energy savings persist across quantization schemes.
  • Investigating replacement of other transformer components — the authors explicitly declined to use LUT layers for the token mixer, so whether any attention-related block benefits remains open.
  • Applying the approach to additional domains and workloads beyond the four datasets studied, and further tuning the LUT configuration parameters (thermometer encoding width, LUT layer count, LUT counts per layer, encoded-value precision), which the paper notes can be scaled to the backbone.

Target Audience

Researchers and practitioners in efficient machine learning, FPGA accelerator design and algorithm-hardware co-design; engineers building edge vision systems where power, latency and on-chip memory are hard constraints; and readers already familiar with vision transformers, quantization, and look-up-table neuron networks who want to see how those ideas combine into an end-to-end deployable system.

Authors’ abstract

Vision Transformers have been tremendously successful in computer vision tasks. However, their large computational, memory, and energy demands are a challenge for edge inference on FPGAs -- a field that has seen a recent surge in demand. We recognize the benefits of recent works on logic and Look Up Table (LUT) based networks, such as LogicNets, NeuraLUT, DWN, among others, in offering models that simultaneously reduce both the memory and compute footprints. However, these models natively do not perform well on common vision tasks, such as CIFAR-10/100. In this work, we propose LL-ViT, a novel edge optimized vision transformer design that integrates layers of LUT neurons within the transformer architecture. Based on our characterization that reveals that a majority of model weights and computations are from the channel mixer (MLP layer), we design an alternate LUT-based channel mixer, and simultaneously develop an FPGA-based accelerator for LL-ViT. Contrary to some attempts to replace each multiplication with a table lookup, our architecture utilizes a neural learning approach which natively learns the LUT functions. This approach allows for reduced model sizes, and a computational and energy-efficient inference solution for vision transformer models. Evaluating on edge-suitable workloads, we achieve accuracies of 95.5% on CIFAR-10, 78.8% on CIFAR-100, and 60.9% on Tiny-ImageNet datasets, comparable to the baseline transformer. LL-ViT eliminates over 60% of the model weights and 50% of the multiplications in the model, and achieves 1.9x energy efficiency and 1.3x lower latency over an integer quantized ViT accelerator, while also offering superior throughput against prior works at a 10.9W power budget.

Read the original paper