NVIDIA TensorRT 7’s Compiler Delivers Real-Time Inference for Smarter Human-to-AI Interactions

TensorRT 7 features a new deep learning compiler designed to automatically optimize and accelerate the complex recurrent and transformer-based neural networks needed for AI speech applications.

Latest Digital Thread News

Latest Digital Thread Resources

Design & Simulation Software Guide 2025

In this Special Issue, Digital Engineering presents its second annual guide to design and simulation software vendors.
Design & Simulation Software Guide

In this Special Issue, Digital Engineering presents its inaugural guide to design and simulation software vendors, including listings for CAD, CAM, simulation, generative design, PLM, rendering and visualization, design for additive manufacturing,…
More Resources

By DE Editors

December 20, 2019

NVIDIA introduced inference software that developers everywhere can use to deliver conversational AI applications.

NVIDIA TensorRT 7—the seventh generation of the company’s inference software development kit—enables smarter human-to-AI interactions, enabling real-time engagement with applications such as voice agents, chatbots and recommendation engines. TensorRT 7 features a new deep learning compiler designed to automatically optimize and accelerate the complex recurrent and transformer-based neural networks needed for AI speech applications.

“We have entered a new chapter in AI, where machines are capable of understanding human language in real time,” said NVIDIA founder and CEO Jensen Huang at his GTC China keynote. “TensorRT 7 helps make this possible, providing developers everywhere with the tools to build and deploy faster, smarter conversational AI services that allow more natural human-to-AI interaction.”

Importance of Recurrent Neural Networks

TensorRT 7 speeds up a growing universe of AI models that are being used to make predictions on time-series, sequence-data scenarios that use recurrent loop structures, called RNNs. In addition to being used for conversational AI speech networks, RNNs help with arrival time planning for cars or satellites, prediction of events in electronic medical records, financial asset forecasting and fraud detection.

With TensorRT’s new deep learning compiler, developers everywhere now have the ability to automatically optimize networks—such as bespoke automatic speech recognition networks, and WaveRNN and Tacotron 2 for text-to-speech—and to deliver the best possible performance and lowest latencies.

The new compiler also optimizes transformer-based models like BERT for natural language processing.

Accelerating Inference from Edge to Cloud

TensorRT 7 can rapidly optimize, validate and deploy a trained neural network for inference by hyperscale data centers, embedded or automotive GPU platforms.

NVIDIA’s inference platform, which includes TensorRT, as well as several NVIDIA CUDA-X AI libraries and NVIDIA GPUs—delivers low-latency, high-throughput inference for applications beyond conversational AI, including image classification, fraud detection, segmentation, object detection and recommendation engines. Its capabilities are widely used by some of the leading enterprise and consumer technology companies, including Alibaba, American Express, Baidu, PayPal, Pinterest, Snap, Tencent and Twitter.

Availability

TensorRT 7 will be available in the coming days for development and deployment, without charge to members of the NVIDIA Developer program from the TensorRT webpage. The latest versions of plug-ins, parsers and samples are also available as open source from the TensorRT GitHub repository.

Sources: Press materials received from the company and additional information gleaned from the company’s website.

More about NVIDIA

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and…

Cut Retrieval-Augmented Generation (RAG) Hallucinations by 50%

Most teams hit the same wall with enterprise AI: LLMs that hallucinate, pipelines that don’t scale, and infrastructure that’s harder to design than the models themselves.

Latest in NVIDIA

About DE Editors

DE's editors contribute news and new product announcements to Digital Engineering. Press releases may be sent to them via [email protected].

Follow DE
on Facebook
on Linkedin

NVIDIA TensorRT 7’s Compiler Delivers Real-Time Inference for Smarter Human-to-AI Interactions

TensorRT 7 features a new deep learning compiler designed to automatically optimize and accelerate the complex recurrent and transformer-based neural networks needed for AI speech applications.

Latest Digital Thread News

Latest Digital Thread Resources

Importance of Recurrent Neural Networks

Accelerating Inference from Edge to Cloud

Availability

More about NVIDIA

Latest in NVIDIA

Latest in NVIDIA

About DE Editors

Related Topics

Subscribe

Subscribe to our FREE magazine, FREE email newsletters or both!

From our Sponsors

Digital Engineering 24/7

Design

Simulate

Additive

Digital Thread

Computing

Resources

Our Partners

Design

Top Story

Latest in Design

Simulation

Top Story

Latest in Simulation

Additive Manufacturing

Top Story

Latest in Additive Manufacturing

Digital Thread

Top Story

Latest in Digital Thread

Engineering Computing

Top Story

Latest in Engineering Computing

Subscribe

Latest Magazine

Latest Special Issue

Newswire

NVIDIA TensorRT 7’s Compiler Delivers Real-Time Inference for Smarter Human-to-AI Interactions

TensorRT 7 features a new deep learning compiler designed to automatically optimize and accelerate the complex recurrent and transformer-based neural networks needed for AI speech applications.

Latest Digital Thread News

Latest Digital Thread Resources

Importance of Recurrent Neural Networks

Accelerating Inference from Edge to Cloud

Availability

More about NVIDIA

Latest in NVIDIA

Latest in NVIDIA

About DE Editors

Related Topics

Subscribe

Subscribe to our FREE magazine, FREE email newsletters or both!

From our Sponsors

Digital Engineering 24/7

Design

Simulate

Additive

Digital Thread

Computing

Resources

Our Partners