At the software level, LightNobel proposes Token-wise Adaptive Activation Quantization (AAQ), which leverages unique token-wise characteristics, such as distogram
patterns in PPM activations, to enable fine-grained quantization techniques without compromising accuracy. At the hardware level, LightNobel integrates the multi-precision reconfigurable matrix processing unit (RMPU) and versatile vector processing unit (VVPU) to enable the efficient execution of AAQ. Through these innovations, LightNobel achieves up to 8.44×, 8.41× speedup and 37.29×, 43.35× higher power efficiency over the latest NVIDIA A100 and H100 GPUs, respectively, while maintaining negligible accuracy loss. It also reduces the peak memory requirement up to 120.05× in PPM, enabling scalable processing for proteins with long sequences
