NVIDIA Brings Confidential AI Inference to Blackwell GPUs
Darius Baruo
Sep 22, 2026 18:34
NVIDIA’s confidential computing delivers secure AI inference on Blackwell GPUs, retaining 96-98% performance while safeguarding sensitive data.
NVIDIA has unveiled its latest advancements in confidential computing, enabling secure AI inference on its Blackwell GPUs while maintaining high performance. According to a company blog post published September 22, 2026, NVIDIA’s Confidential Computing (CC) framework now supports the execution of large language model (LLM) workloads in memory-encrypted environments, retaining 96-98% of the performance seen in non-confidential configurations.
Confidential computing is designed to protect sensitive data and proprietary AI models during active processing, expanding on NVIDIA’s earlier Hopper-generation H100 GPU capabilities. This solution caters to high-value, regulated industries like healthcare, financial services, and government, where data security during inference is critical. By incorporating confidential virtual machines (CVMs), encrypted NVLink, and confidential Blackwell GPUs, NVIDIA aims to make secure deployments viable without significant performance trade-offs.
Performance Benchmarked: Minimal Overhead
In its tests, NVIDIA evaluated the performance impact of enabling confidential computing using its TensorRT LLM inference framework. The comparison involved running identical workloads under two conditions: with CC enabled and disabled. Workloads were processed on NVIDIA’s DGX B200 systems, featuring eight Blackwell GPUs, using Intel TDX CVMs for secure execution.
Key metrics included output-token throughput and time-per-output-token (TPOT) latency. At concurrency levels of 1-16, CC-enabled inference retained 96.1-98.2% of the baseline throughput and introduced a latency overhead of just 1.2-4.3%. For organizations requiring strict security measures, these benchmarks demonstrate that confidentiality can coexist with near-peak performance.
Adapting AI Frameworks for Security
Implementing confidential computing requires adaptations at the framework level. For instance, NVIDIA adjusted TensorRT LLM’s data movement mechanisms to account for encrypted host-to-device transfers, which bypass direct GPU access to protected memory. NVIDIA also optimized multi-GPU communication, compensating for the absence of NVLink SHARP multicast in CC configurations.
These adaptations ensure secure execution while minimizing disruptions to model inference pipelines. The company emphasized that such configurations should be treated as a full-stack problem, requiring engineers to optimize both security and performance concurrently.
Strategic Implications
With the market for AI hardware expected to surpass $150 billion by 2030, per industry estimates, NVIDIA’s push into confidential computing aligns with growing demand for secure AI solutions. The technology has immediate applications in sectors where trust, privacy, and regulatory compliance are paramount. For example, healthcare providers can securely analyze sensitive patient data, while financial institutions can execute proprietary models without risking intellectual property exposure.
For traders, NVIDIA’s continued innovation strengthens its dominance in the AI hardware market. As of September 22, NVIDIA’s stock price stands at $229.26, up 0.82% over the past 24 hours, with a market cap of $5.567 trillion. Investors betting on AI-driven growth should monitor the adoption of confidential computing in enterprise and government deployments, as this could further solidify NVIDIA’s leadership in AI and GPU-accelerated computing.
To explore NVIDIA’s confidential computing offerings and start planning secure AI deployments, developers can access the NVIDIA Trusted Computing documentation and the latest TensorRT LLM features. With its Blackwell GPUs now setting the benchmark for secure, high-performance AI inference, NVIDIA is positioning itself as the go-to provider for confidential AI infrastructure.
Image source: Shutterstock
