Free Lesson
Optimizing Search & Data Processing with Self-hosted SLMs
Part of Exploring Modern AI Search
60 min
Feb 27, 2026 11:00 AM
Virtual (Zoom)
In this video
What you'll learn
Search & Data Processing with Small Language Models
See what tasks are ideal for switching from LLMs to SLMs with the same or better quality.
Architecting an inference engine
Lessons from wrapping SGLang, vLLM, TensorRT and why we chose to rewrite most of the OSS model code.
Support 100s of task-specific models
How to support 35+ model architectures without going crazy and how to run LoRAs in production.
1 million tokens per second in an OSS K8s cluster
Design an auto-scaled multi-model multi-modal cluster to drive all GenAI projects in your business.
Why this topic matters
In this Lightning Lesson, Daniel will share a preview of Superlinked Inference Engine - an open source software for self-hosting Small Language Models in your own cloud. Cut 95%+ of your managed LLM API cost, gain access to a wide catalog of esoteric and fine-tunable SOTA models, and regain security and control by self-hosting.
You'll learn from

Daniel Svonava
CEO at Superlinked

Trey Grainger
Founder at Searchkernel, Author "AI-Powered Search".

Doug Turnbull
Principal Search Consultant
Previously at
