EP02

Unlocking AI Potential: Exploring Offline LLMs

Johan van Amersfoort joins us to talk about offline and private large language models, GPUs, Kubernetes, model choices, and his chatbot project Koby.

With Johan van Amersfoort, Chief Evangelist, ITQ

Introduction

In this episode, we talk with Johan van Amersfoort about offline and private large language models, GPUs, Kubernetes, model choices, and the chatbot project Koby.

Meet the Guest

Johan van Amersfoort is Chief Evangelist at ITQ. He works within the CTO office on new technology and has a background in digital workspace, virtualization and modern application platforms.

Setting the Stage

Public AI services are easy to use, but not always suitable for sensitive data. In healthcare, government and other regulated sectors, it can be necessary to run models locally or within your own environment.

Episode Highlights

  • Johan built Koby to find out whether a chatbot could stand in for him during a long vacation.
  • Factual VMware documentation and his own, more opinionated books produced different answers.

Deep Dive

A local model is only one part of a private AI solution. You also need runtime, storage, frameworks, GPU or CPU capacity, security and monitoring. Johan used tools including Mistral, NVIDIA NeMo, Hugging Face and Kubernetes. Models and containers can be large. For public data, a cloud chatbot can work fine; for sensitive data, an on-premises solution may be a better fit.

Real-Life Stories & Examples

  • Koby was fed VMware documentation, ITQ information and Johan's own books.
  • Guardrails are needed to keep a model within its intended purpose and approved topics.
  • An early version used a public API combined with Johan's own information via RAG.

Key Takeaways

  • Offline AI can be well suited to sensitive and regulated data.
  • Private AI requires hardware, software, management and expertise.
  • Models inherit bias from their source data.
  • Guardrails limit unwanted answers.
  • The right architecture depends on data, workload, budget and compliance.

Closing Thoughts

Offline LLMs make AI more controllable, but not automatically simpler. Start with a concrete problem, build something small, and only then figure out which architecture you actually need.