
NVIDIA and AWS push mid-range AI inference with Blackwell hardware and GPU-accelerated vector search
Published by AINave Editorial • Reviewed by Ramit
Nvidia and AWS have expanded their partnership to make AI inference and vector search more accessible for production workloads. The new EC2 G7 instance brings Blackwell GPUs to a mid-tier price point, while OpenSearch Serverless now uses Nvidia's cuVS library for GPU-accelerated vector indexing by default. For AI builders, this means faster inference and retrieval without the operational overhead of managing custom GPU pipelines.
What happened
On June 24, 2026, Nvidia detailed two major updates to its AWS collaboration. The first is the EC2 G7 instance, powered by Nvidia's RTX PRO 4500 Blackwell Server Edition GPUs paired with custom sixth-generation Intel Xeon processors. AWS says it delivers up to 4.6 times the AI inference performance and 2.1 times the graphics performance of the previous G6 generation. The instance supports up to eight GPUs with 256GB of total GPU memory, 700 Gbps of EFA networking (seven times G6), and up to 7.6TB of local NVMe storage. It reached general availability on June 18, 2026, in AWS's Ohio and Oregon regions.
The second update is on the retrieval side. AWS made Nvidia's cuVS library the default for vector indexing across all collections in next-generation OpenSearch Serverless. Nvidia says cuVS makes vector indexing up to 10 times faster than CPU-only builds at a quarter of the cost, and can build billion-scale vector databases in under an hour. This turns GPU-accelerated vector search from a specialized optimization project into a native AWS capability.
Why AI builders should care
The G7 is positioned as a mid-tier option for inference, rendering, video, spatial computing, and virtual desktops. It sits below the G7e family (launched in January 2026 on the larger RTX PRO 6000 Blackwell GPU). If G7e is the premium tier for bigger generative models, G7 is the volume tier that puts Blackwell into everyday inference and analytics jobs without committing teams to top-of-rack hardware they don't need.
For teams building retrieval-augmented generation (RAG), semantic search, recommendation systems, or agentic AI, the cuVS integration in OpenSearch Serverless removes a major operational hurdle. Developers no longer need to stand up and tune their own GPU pipeline for vector search. The optimization becomes a checkbox, not a project.
Practical implications
Enterprises can now deploy mid-tier GPU-accelerated inference workloads with less upfront cost and maintenance than high-end training clusters. The G7 instance is designed for the "messy middle" of enterprise AI: the compute that serves a model and the retrieval that feeds it. Combined with native GPU-accelerated vector search, teams can build production AI pipelines faster and at lower cost.
AWS also gains a competitive edge. Google Cloud and Microsoft Azure have not yet shipped a comparable Blackwell-accelerated managed instance for the mid-range. If the next phase of cloud AI is decided by who makes the right-sized GPU easy and cheap to run, this announcement targets that contest directly.
Caveats
Performance claims (4.6x AI inference, 2.1x graphics) are AWS-reported versus the G6 generation
Sources
- Nvidia and AWS Deepen Ties to Speed AI Inference and Vector Search
- NVIDIA and AWS Expand Full-Stack Partnership
- Can AI infrastructure be both sovereign AND scalable? | Helen Yu
- Nvidia and AWS Deepen AI Partnership for Enterprise Scale
- Nvidia and AWS expand partnership with specialized AI hardware and ...
- Nvidia and AWS deepen push to simplify AI infrastructure at scale
- Cerebras After the IPO: OpenAI, AWS and the Fight for AI Inference
- AWS Summit NY 2026: Is AI Infrastructure AWS's Real Agentic ...
- Nvidia and AWS deepen push to simplify AI infrastructure at scale
- NVIDIA GPU-Accelerated Amazon Web Services
- Nvidia and AWS GTC 2026 Push Production AI Control Plane, Not ...



















