Vivek Singh

Senior Cloud Engineer
dunnhumby
India

About

Vivek Singh is a Senior Cloud Engineer specializing in cloud architecture, platform engineering, and data solutions using Microsoft Azure and modern cloud-native technologies. He has extensive experience designing and delivering secure, scalable cloud platforms, infrastructure as code (IaC) using Terraform, Azure Kubernetes Service (AKS), GitOps, CI/CD automation, cloud governance, and AI-enabled developer platforms.A Microsoft Certified Trainer (MCT) since 2017, Vivek has delivered numerous technical workshops and conference sessions at national and international events, helping engineers and organizations adopt Azure, cloud-native architectures, DevOps, and modern data engineering practices. His work focuses on building enterprise-scale cloud platforms, enabling AI-driven developer productivity, and implementing real-time analytics and data solutions using Azure and Databricks.He is passionate about sharing practical knowledge with the community and enjoys speaking about cloud infrastructure, platform engineering, AI, and data technologies.
Talk

Stop Paying for Unused Tokens: AI Cost Optimization with Vertex AI

Generative AI, LLM Cost Optimization, Vertex AI, Prompt Engineering
As enterprises rapidly adopt large language models (LLMs), managing inference costs has become a critical challenge. Many AI applications repeatedly transmit large amounts of static or redundant context, resulting in unnecessary token consumption, increased latency, and higher operational expenses. While organizations invest heavily in AI capabilities, few optimize how prompts are constructed and delivered to the model.In this session, Vivek Singh demonstrates practical techniques for reducing LLM token usage by up to 50% through intelligent prompt preprocessing, semantic context selection, and context caching in Google Vertex AI. Attendees will learn how to eliminate redundant information, preprocess structured and unstructured data, retrieve only the most relevant context, and leverage Vertex AI’s context caching capabilities to avoid repeatedly sending identical information.Through a live demonstration, the audience will compare a traditional “send everything” approach with an optimized pipeline, examining token consumption, response latency, and API costs in real time. The session also covers implementation patterns, architectural best practices, and measurable strategies that can be integrated into enterprise AI applications without sacrificing response quality.Attendees will leave with the practical knowledge needed to build faster, more scalable, and significantly more cost-efficient generative AI solutions, making this session valuable for AI engineers, data engineers, cloud architects, platform engineers, and developers deploying LLM-powered applications in production.