NVIDIA’s developer blog discusses co-designing AI model attention mechanisms to improve speed for long-context, interactive workloads, noting that as such workloads grow, attention can become a larger portion of inference time.
As agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1). Because...