Baseten / Blog

Baseten Blog | Page 1

Topics

Latest Model performance Hacks & projects GPU guides ML models Glossary Community Product News

1 2 3…10

Product

May 1, 2024

New in April 2024

Use four new best in class LLMs, stream synthesized speech with XTTS, and deploy models with CI/CD

Baseten

Prompt: the steps and entrance to a solarpunk museum

Hacks & projects

Apr 30, 2024

CI/CD for AI model deployments

In this article, we outline a continuous integration and continuous deployment (CI/CD) pipeline for using AI models in production.

Vlad Shulman

Samiksha Pal

2 others

Hacks & projects

Apr 18, 2024

Streaming real-time text to speech with XTTS V2

In this tutorial, we'll build a streaming endpoint for the XTTS V2 text to speech model with real-time narration and 200 ms time to first chunk.

Het Trivedi

Philip Kiely

Prompt: A wooden boat full of books floating down a rapid river in a Japanese garden

Glossary

Apr 5, 2024

Continuous vs dynamic batching for AI inference

Learn how to increase throughput with minimal impact on latency during model inference with continuous and dynamic batching.

Matt Howard

Philip Kiely

Prompt: A batch of candy being processed on a fantasy assembly line

Product

Mar 28, 2024

New in March 2024

Fast Mistral 7B, fractional H100 GPUs, FP8 quantization, and API endpoints for model management.

Baseten

GPU guides

Mar 28, 2024

Using fractional H100 GPUs for efficient model serving

Multi-Instance GPUs enable splitting a single H100 GPU across two model serving instances for performance that matches or beats an A100 GPU at a 20% lower cost.

Matt Howard

Vlad Shulman

2 others

Prompt: Two tron-style motorcycles racing on an empty highway

Model performance

Mar 14, 2024

Benchmarking fast Mistral 7B inference

Running Mistral 7B in FP8 on H100 GPUs with TensorRT-LLM, we achieve best in class time to first token and tokens per second on independent benchmarks.

Abu Qader

Pankaj Gupta

2 others

Prompt: a model bullet train in a snowy village.

Model performance

Mar 14, 2024

33% faster LLM inference with FP8 quantization

Quantizing Mistral 7B to FP8 resulted in near-zero perplexity gains and yielded material performance improvements across latency, throughput, and cost.

Pankaj Gupta

Philip Kiely

Prompt: A ship in a bottle in a dark wood library

Model performance

Mar 12, 2024

High performance ML inference with NVIDIA TensorRT

Use TensorRT to achieve 40% lower latency for SDXL and sub-200ms time to first token for Mixtral 8x7B on A100 and H100 GPUs.

Justin Yi

Philip Kiely

Prompt: A friendly robot horse playing in a sunlit meadow

Glossary

Mar 7, 2024

FP8: Efficient model inference with 8-bit floating point numbers

The FP8 data format has an expanded dynamic range versus INT8 which allows for quantizing weights and activations for more LLMs without loss of output quality.

Pankaj Gupta

Philip Kiely

1 2 3…10