Get Trained and Certified at GTC Paris at VivaTech 2025

SOURCE | 1 year ago


Enhance your Social Media content with NViNiO•AI™ for FREE


Retrieval-Augmented Generation (RAG) pipelines are revolutionizing enterprise operations. However, most existing tutorials stop at proof-of-concept implementations that falter when scaling. This workshop aims to bridge that gap, focusing on building scalable, production-ready RAG pipelines powered by NVIDIA NIM microservices and Kubernetes. Participants will gain hands-on experience deploying, monitoring, and scaling RAG pipelines with the NIM Operator and learn best practices for infrastructure optimization, performance monitoring, and handling high traffic volumes.

The workshop begins by building a simple RAG pipeline using the NVIDIA API catalog. Participants will deploy and test individual components in a local environment using Docker Compose. Once familiar with the basics, the focus will shift to deploying NIMs, such as LLM, NeMo Retriever Text Embedding, and NeMo Retriever Text Reranking, in a Kubernetes cluster using the NIM Operator. This'll include managing the deployment, monitoring, and scalability of NVIDIA NIM microservices. The workshop will focus on building a RAG pipeline that can be used in production. It'll also look at the NVIDIA AI Blueprint for PDF ingestion, and how to use it in the RAG pipeline.

To ensure operational efficiency, the workshop will introduce Prometheus and Grafana for monitoring pipeline performance, cluster health, and resource utilization. Scalability will be addressed through the use of the Kubernetes Horizontal Pod Autoscaler (HPA) for dynamically scaling NIMs based on custom metrics in conjunction with the NIM Operator. Custom dashboards will be created to visualize key metrics and interpret performance insights.

You'll be able to:

Build a simple RAG pipeline using API endpoints, deployed locally with Docker Compose. Deploy a variety of NVIDIA NIM microservices in a Kubernetes cluster using the NIM Operator. Combine NIMs into a cohesive, production-grade RAG pipeline and integrate advanced data ingestion workflows. Monitor RAG pipelines and Kubernetes clusters with Prometheus and Grafana. Scale NIMs to handle high traffic using the NIM Operator. Create, deploy, and scale RAG pipelines for a variety of agentic workflows, including PDF ingestion.

Prerequisite(s)

Familiarity working with LLM-based applications Familiarity with RAG pipelines Familiarity working with Kubernetes Familiarity working with Helm

Enhance your brand's digital communication with NViNiO•Link™ : Get started for FREE here


Read Entire Article

© 2026 | Actualités Africaines & Tech | Moteur de recherche. NViNiO GROUP

_