Generative Adversarial Networks

Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, Yoshua Bengio

arXiv (Cornell University) · 2014 · 인용 4.6k

Large Language Models (LLMS) rely on Key-Value (KV) caches to store attention context during autoregressive decoding. In long-sequence settings, the KV cache can consume large amounts of VRAM and become a practical bottleneck for throughput . We introduce KVHALO, an auxiliary reconstruction model that restores higher-fidelity KV tensors from a compressed cache state when required, reducing persistent memory footprint during inference.

In our evaluation, KVHALO achieves up to 91.85% directional cosine alignment at convergence and reduces long-context degradation relative to a low-bit baseline under our stress-test workloads. We used HRM instead of other architectures, which allowed for higher-quality results in only 18,600 steps.

🏛️ 거인의 어깨이 분야를 만든 논문들

생성적 적대 신경망(GAN)을 제안하여 데이터 생성 및 변환 분야에서 인공지능의 새로운 가능성을 열었습니다.

이야기를 쓰는 중…

Paperis - Generative Adversarial Networks