Amazon Web Services (AWS) today announced the general availability of Amazon Elastic Compute Cloud (Amazon EC2) G7 instances, marking a significant leap forward in cloud-based GPU acceleration for a wide array of demanding workloads. This release positions AWS as the first major cloud provider to offer instances powered by the NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs, paired with custom sixth-generation Intel Xeon Scalable processors. The new G7 instances are engineered to deliver substantial performance improvements, offering up to 4.6 times faster AI inference and up to 2.1 times greater graphics performance compared to the previous generation G6 instances. This advancement is set to revolutionize capabilities for artificial intelligence (AI) inference, high-fidelity graphics rendering, and accelerated data analytics, particularly for workloads running on Amazon EMR and Amazon Elastic Kubernetes Service (Amazon EKS).
The introduction of G7 instances underscores AWS’s commitment to providing cutting-edge infrastructure that meets the escalating demands of modern, compute-intensive applications. From the foundational elements of machine learning model deployment to the intricate processes of virtual production and scientific simulation, the need for robust, scalable GPU resources has never been more critical. The G7 instances are specifically tailored for a broad spectrum of GPU-enabled applications, including generative AI inference, sophisticated graphics rendering, high-throughput video transcoding and analytics, immersive spatial computing environments, high-performance virtual desktop infrastructure (VDI), and advanced data analytics. This versatility ensures that a diverse range of industries, from media and entertainment to scientific research and financial services, can leverage the power of the cloud to accelerate their most complex tasks.
A New Era of Cloud GPU Performance with NVIDIA Blackwell
The core of the G7 instance’s groundbreaking performance lies in its integration of the NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs. NVIDIA’s Blackwell architecture, unveiled earlier this year, represents a monumental stride in GPU design, succeeding the highly successful Hopper architecture. While the broader Blackwell platform targets large-scale AI training and supercomputing, the RTX PRO 4500 Server Edition is specifically optimized for enterprise-grade professional visualization, AI development, and data processing. These GPUs are designed to handle demanding graphics workloads with unparalleled efficiency, featuring enhanced ray tracing capabilities, advanced Tensor Cores for AI acceleration, and significant improvements in memory bandwidth and processing power.
AWS’s early adoption of Blackwell technology in a general-purpose cloud offering is a testament to its strategic partnership with NVIDIA, a collaboration that has consistently pushed the boundaries of what’s possible in cloud computing. By being the first major cloud provider to support these new GPUs, AWS is providing its customers with a distinct advantage, enabling them to innovate faster and achieve previously unattainable performance benchmarks for their most critical applications. The Blackwell architecture is characterized by its modular design, allowing for flexible configurations that can scale from individual workstations to massive data centers, making it an ideal fit for the elastic and scalable nature of AWS EC2.
Complementing the NVIDIA GPUs are custom sixth-generation Intel Xeon Scalable processors. This synergy between leading GPU and CPU technology ensures that G7 instances are not only powerful in their graphical and AI processing capabilities but also robust in general-purpose computing. Intel’s latest Xeon processors bring advancements in core density, memory bandwidth, and I/O throughput, providing a balanced and high-performance foundation for the GPUs to operate at their peak. This custom integration is crucial for minimizing bottlenecks and maximizing the overall efficiency of the instance, ensuring that data can be fed to and processed by the GPUs at an optimal rate.
Unpacking the Performance Gains and Use Cases
The headline performance figures—up to 4.6 times faster AI inference and up to 2.1 times greater graphics performance compared to G6 instances—translate into substantial real-world benefits. For AI inference, particularly with the proliferation of generative AI models and large language models (LLMs), faster processing means quicker response times, lower latency for real-time applications, and the ability to handle higher throughput of requests. This is critical for applications like conversational AI, content generation, personalized recommendations, and sophisticated computer vision systems where immediate insights are paramount. A 4.6x improvement can drastically reduce operational costs and enhance user experience for AI-powered services.
In the realm of graphics, the 2.1x performance increase is transformative for industries reliant on visual fidelity and complex simulations. This includes virtual production studios, architectural visualization firms, game development, and engineering design. High-resolution rendering, real-time ray tracing, and intricate simulations can now be performed with unprecedented speed and detail in the cloud, empowering creators and engineers to iterate faster and bring more ambitious projects to life. For virtual desktop infrastructure (VDI), this means supporting more demanding professional applications like CAD/CAM, video editing, and 3D modeling with desktop-like responsiveness, even for remote users.
Beyond AI and graphics, G7 instances also deliver faster performance for GPU-accelerated analytics. When integrated with services like Amazon EMR on Amazon EKS, these instances can significantly accelerate the processing of large datasets, enabling quicker insights from complex analytical workloads. This is particularly beneficial for tasks such as financial modeling, scientific data analysis, and large-scale data warehousing, where the parallel processing power of GPUs can dramatically reduce computation times. The ability to deploy these accelerated analytics within a Kubernetes environment through EKS offers developers and data scientists flexibility and scalability for their containerized applications.
Technical Specifications and Advanced Networking Capabilities
The G7 instance family is designed with scalability and performance at its core, offering a range of configurations to suit various workload requirements. These instances feature up to 8 NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs, each equipped with 32 GB of memory, culminating in a formidable 256 GB of total GPU memory for the largest configurations. This substantial GPU memory is crucial for handling large AI models and high-resolution graphical assets.

The instances are available in seven distinct sizes, with an additional g7.metal bare-metal option slated for future release. The largest virtualized instance, g7.48xlarge, boasts 192 vCPUs, up to 768 GiB of system memory, and up to 7.6 TB of local NVMe SSD storage. Network bandwidth is a critical component of high-performance computing, and G7 instances excel here, offering up to an astonishing 700 Gbps of network bandwidth. This massive throughput is vital for data-intensive applications, ensuring that GPUs are not starved for data and can process information with minimal latency. EBS bandwidth is also robust, reaching up to 80 Gbps, facilitating rapid access to persistent storage.
A key differentiator for G7 instances, especially for multi-GPU and multi-node workloads, is their support for NVIDIA GPUDirect P2P, GPUDirect RDMA with EFA, and GPUDirect RDMA with EFA for Amazon FSx for Lustre. These technologies are foundational for high-performance computing (HPC) environments. GPUDirect P2P enables direct communication between GPUs within the same instance, bypassing the CPU and system memory to significantly reduce latency. GPUDirect RDMA (Remote Direct Memory Access) with Elastic Fabric Adapter (EFA) extends this capability across multiple instances, allowing GPUs in different servers to communicate directly without CPU involvement, which is paramount for scaling AI training and HPC simulations across large clusters. Furthermore, integration with Amazon FSx for Lustre, a high-performance file system, ensures that data can be accessed and processed by GPUs at extreme speeds, eliminating I/O bottlenecks that often plague large-scale computing tasks.
Ecosystem Integration and Ease of Adoption
AWS has ensured that G7 instances are easily accessible and integrated into its comprehensive ecosystem. To facilitate rapid deployment for AI inference and graphics workloads, customers can leverage the AWS Deep Learning AMIs (DLAMI) or NVIDIA Workstation AMIs, which come pre-packaged with the necessary GPU drivers and optimized software stacks. For containerized applications orchestrated with Amazon EKS, users can build EKS AMIs using NVIDIA driver version R595 with EKS-provided automation, ensuring seamless compatibility and performance.
The G7 instances also support a wide range of popular operating systems, including Amazon Linux, Ubuntu, Red Hat Enterprise Linux (RHEL), and Windows Server. This broad OS compatibility, coupled with comprehensive NVIDIA driver integration, ensures support for industry-standard graphics libraries such as DirectX, Vulkan, and OpenGL. This makes G7 instances ideal for developers and enterprises working with existing graphics-intensive applications or developing new ones across various platforms. The flexibility in OS choice and driver support lowers the barrier to entry for migrating diverse workloads to the cloud.
Availability and Economic Considerations
The Amazon EC2 G7 instances are currently available in two key AWS regions: US East (Ohio) and US West (Oregon). AWS typically rolls out new instance types to additional regions based on customer demand and infrastructure readiness, with customers able to monitor future regional expansion plans via the CloudFormation resources tab on the AWS Capabilities by Region page. This phased rollout strategy ensures stability and optimal performance from the outset.
AWS offers multiple flexible purchasing options for G7 instances, catering to various financial strategies and workload patterns. Customers can choose On-Demand pricing for immediate access and pay-as-you-go flexibility, suitable for experimental or fluctuating workloads. For predictable, long-running commitments, Savings Plans provide significant discounts compared to On-Demand rates. Spot Instances offer an even more cost-effective solution for fault-tolerant workloads, leveraging unused EC2 capacity for substantial savings. Additionally, Dedicated Instances are supported for the larger 12xlarge, 24xlarge, and 48xlarge sizes, providing dedicated physical servers for customers requiring strict resource isolation and compliance. These diverse purchasing models empower businesses to optimize their cloud spending while accessing state-of-the-art GPU compute.
Broader Impact and Implications
The launch of G7 instances represents more than just a new set of hardware; it signifies a strategic move by AWS to solidify its leadership in the rapidly expanding market for AI and high-performance computing infrastructure. As AI adoption accelerates across all sectors, the demand for specialized, high-performance GPUs in the cloud will only intensify. By being an early mover with NVIDIA’s Blackwell architecture, AWS is positioning itself at the forefront of this evolution, enabling its customers to build and deploy more sophisticated AI models and immersive graphics experiences than ever before.
The implications for various industries are profound. In healthcare, G7 instances could accelerate drug discovery through faster molecular simulations and enhance medical imaging analysis with advanced AI. In manufacturing, they can power digital twins with greater fidelity and enable real-time simulation for product design and factory optimization. The media and entertainment industry will benefit from faster rendering times for visual effects, animation, and virtual production, drastically reducing production cycles. Even in financial services, complex algorithmic trading and risk analysis models can leverage the G7’s speed for quicker decision-making and deeper insights.
This release also intensifies the competitive landscape among major cloud providers. While other providers are also investing heavily in GPU infrastructure, AWS’s claim of being the first to offer Blackwell-powered instances provides a temporary but significant lead, potentially attracting enterprises and startups that require the absolute latest in GPU technology. The continuous innovation in cloud GPU offerings is a clear indicator of the industry’s focus on democratizing access to high-performance computing, making capabilities once reserved for supercomputers available to a broader audience on a pay-as-you-go basis.
In conclusion, the general availability of Amazon EC2 G7 instances marks a pivotal moment in cloud computing. By integrating NVIDIA’s cutting-edge Blackwell GPUs with custom Intel Xeon processors and AWS’s robust infrastructure, these instances offer unprecedented performance for AI inference, graphics, and data analytics. This development empowers businesses across diverse sectors to accelerate their innovation, reduce operational costs, and push the boundaries of what is possible with cloud-based high-performance computing. AWS continues to solicit customer feedback on these new instances, inviting users to share their experiences on AWS re:Post for EC2 or through their usual AWS Support contacts, ensuring that future iterations of their services remain aligned with the evolving needs of their global customer base.
