Series: AI Workloads for Database Engineers | Part 1 of 6
Introduction
In recent years, many organizations have started adding AI capabilities to their applications. Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) and Agentic AI are being used for semantic search, document assistants, recommendations and business workflows.
For database engineers, this often starts with something that looks quite simple. An embedding model converts content into vectors, which are then used for similarity search. In a common setup, the application team adds an embedding column to an existing table, stores vectors alongside business data and introduces similarity queries. Depending on the database and architecture, those embeddings may be generated outside the database or generated automatically as part of the database's vector search capabilities.
It looks like a small change to the database. But is it really just another column and another query?
Traditional database workloads are built around a familiar pattern. An application knows what it is looking for. The database finds the matching rows using indexes, joins and filters. AI retrieval changes that pattern. The application may not know the exact row it wants. Instead, it asks the database to find rows that are similar to a query.
That difference is what makes AI workloads interesting for DBAs.
Why AI workloads are different?
The key difference is simple:
- Traditional CRUD: exact match.
- AI retrieval: similarity ranking.
A traditional request might look for a customer ID, order number or session token. The database finds the matching row. An AI request sends a vector and asks the database to find the closest matching rows.
That changes the work the database has to do, particularly around CPU, memory, I/O, storage and indexing.
What actually reaches the database?
Three things appear frequently in AI retrieval workloads:
- Embeddings: Fixed-length arrays of numbers generated by a machine learning model. Similar content produces vectors that are close together.
- Vector search: Finding rows closest to a query vector. This can be done through exact search or an approximate vector index.
- Metadata filtering: Vector search is usually combined with normal conditions such as tenant, status or date.
In a typical RAG workflow, the application embeds the user's question, retrieves relevant content from the database and sends that content to an LLM. Agentic AI applications can use the same retrieval pattern as one step in a larger workflow.
Retrieval Vs Ingestion
AI workloads still have the familiar database pattern of writes and reads. The difference is what those operations make the database do.
- Ingestion: Writes embeddings into the database and may require vector index maintenance.
- Retrieval: Searches for similar vectors, ranks the results and often applies additional metadata filters.
These operations can happen at the same time. For example, a large re-embedding job may be updating vectors and maintaining the vector index while users are running similarity searches against the same data.
That is where the DBA starts to care about workload interaction, not just vector storage. The database is still handling reads and writes, but the work involved in those operations can be quite different from a traditional CRUD workload.
What changes for the DBA?
The resources are familiar, but the workload is different.
- CPU: Similarity calculations can require more computation than simple lookups.
- Memory: Vector search and its indexes benefit from sufficient memory.
- I/O: Vector data and indexes can be much larger than traditional transactional indexes.
- Storage: Embeddings and their indexes can add substantial space to tables, backups and replicas.
- Connections: Application code may hold a database connection while waiting on an external AI service.
The actual impact depends on the data size, vector dimensions, index, query pattern, concurrency and configuration.
Wrapping up
AI does not fundamentally change what a database is. It changes the workload being placed on it. The important shift is from finding data that matches to finding data that is similar. That change affects how the database uses CPU, memory, I/O, storage and connections.
For database engineers, understanding this workload is the first step toward making the right decisions around performance and operations.
Series: AI workloads for database engineers
- Part 1: AI workloads are different: What changes for Database Engineers(DBAs) ?
- Part 2: Where should AI data live? (coming soon)
- Part 3: Vector search across PostgreSQL, MongoDB, MySQL and MariaDB (coming soon)
- Part 4: What happens when AI queries hit production? (coming soon)
- Part 5: From Vector Search to RAG: Where Does the Database Fit? (coming soon)
- Part 6: Monitoring and designing production AI database workloads (coming soon)

No comments:
Post a Comment