ScyllaDB Beyond Cassandra, Part 2: Search, CDC and Storage

In ScyllaDB Beyond Cassandra: Tablets, Raft and More, I covered the parts of ScyllaDB that make “Cassandra rewritten in C++” an increasingly incomplete description.
There is another set of differences closer to the application. How do you query a column outside the primary key? Keep a search index current? Run a backfill without overwhelming interactive traffic? How much disk space does compaction need, and do SSTables have to live on local disks at all?
ScyllaDB and Cassandra have taken different approaches to these questions. Some are established features. Others arrived with ScyllaDB 2026.3. They affect both the queries you can write and the infrastructure behind them.
Secondary indexes: similar CQL, different work
Suppose a support application stores tickets by ID. Fetching one ticket is straightforward. Finding tickets by an external reference requires another access path.
Example : find tickets linked to CRM-1842.
CREATE INDEX tickets_external_ref_idx ON support.tickets (external_ref);
SELECT ticket_id, body FROM support.tickets
WHERE external_ref = 'CRM-1842';ScyllaDB's global secondary indexes use materialized views internally. The indexed value becomes the partition key of an index table. A query first reads that table to find the relevant base-table keys, then fetches the rows. Maintaining the index adds work when indexed values change, and the index and base rows can live on different nodes.
Cassandra 5's Storage-Attached Indexing, or SAI, keeps index structures alongside memtables and SSTables on each replica. Queries use those local indexes on the replicas they contact. There is no equivalent global index table partitioned by the indexed value. ScyllaDB's closer counterpart is its local secondary index, which keeps index and base data together but only serves queries that also restrict the partition key.
That affects where reads go, what writes must maintain, and how costs change with selectivity. A query returning ten rows deserves a different test from one matching half the table.
There is also a compatibility detail worth knowing: ScyllaDB accepts Cassandra's StorageAttachedIndex class name for vector indexes, translating it to its own implementation. Accepting that declaration does not mean ScyllaDB runs Cassandra's SAI engine.
Vector search adds another kind of lookup
A ticket ID finds one record. An embedding can help find earlier reports of the same problem, even when the wording differs. “Connection fails after renewing the certificate” might be relevant to “TLS handshake errors following certificate rotation.”
Cassandra already supports vector search, as I covered in the Cassandra 5 and 6 article. ScyllaDB's implementation has a different deployment model.
Its Vector and Text Search service is currently a ScyllaDB Cloud feature. Vector search requires version 2025.4.3 or later, while full-text search requires 2026.3.0 or later. Both require the indexing service to be enabled.
The architecture separates storage nodes from indexing nodes. Rows and their embeddings are stored in ScyllaDB. Separate nodes maintain in-memory HNSW indexes, using USearch for approximate nearest-neighbour search. A CQL query reaches the indexing service, which identifies candidate row keys. Using these keys, ScyllaDB then retrieves the corresponding rows. Indexing capacity can scale separately from database capacity.

Rows and embeddings live on ScyllaDB nodes, while separate indexing nodes maintain vector and full-text indexes through asynchronous CDC updates.
This adds a resource to size. A database may have plenty of storage while its vector index is running out of memory. Vector dimensions, graph configuration and the number of indexed rows all matter. The approximate search also introduces a recall/latency trade-off: returning quickly and finding every relevant neighbour are different objectives.
The application still generates embeddings. It chooses the model, decides how to split long ticket discussions into searchable units, and handles re-embedding when the model changes.
Example: find five similar tickets.
Let’s assume an embedding is stored with each ticket. This value is calculated by an external model, and each model has a fixed output size. With a model that returns 384 values, such as all-MiniLM-L6-v2, the table gets a matching vector column and an index on it:
ALTER TABLE support.tickets ADD embedding vector<float, 384>;
CREATE CUSTOM INDEX tickets_embedding_idx ON support.tickets (embedding)
USING 'vector_index'
WITH OPTIONS = {'similarity_function': 'COSINE'};When a new ticket arrives ("TLS handshake errors following certificate rotation"), the application sends its text to the same model, either a hosted API or a locally run one, and gets back 384 numbers. It passes them to the query as a bind parameter. The query syntax is the same as in Cassandra 5:
SELECT ticket_id, body FROM support.tickets
ORDER BY embedding ANN OF ?
LIMIT 5;Full-text search covers the words you actually typed
Sometimes the useful search term is simply SSLHandshakeException. An exception name, product identifier or quoted phrase can be more precise than a semantic description.
ScyllaDB 2026.3 adds full-text search through fulltext_index indexes and BM25 ranking. For a ticket’s body column:
Example: find tickets mentioning SSLHandshakeException.
CREATE CUSTOM INDEX tickets_body_fts ON support.tickets (body)
USING 'fulltext_index';
SELECT ticket_id, body FROM support.tickets
WHERE BM25(body, 'SSLHandshakeException') > 0
ORDER BY BM25(body, 'SSLHandshakeException')
LIMIT 10;An analyzer splits and normalizes text into terms. The index finds matching documents and ranks them. Language-specific analyzers also handle word forms through stemming. The table must use tablets, and partition-key columns cannot be indexed. Creating a full-text index automatically enables CDC on the base table, adding change-log storage to account for.
Queries require matching BM25() expressions in WHERE and ORDER BY, plus a LIMIT. Paging, grouping and aggregation are currently unsupported.
Expiration needs separate attention. Standard cell-level TTL does not generate CDC events, so the index can retain entries after the underlying values expire. The documentation recommends per-row TTL when the index must track expiration: it generates deletion events.
Full-text indexes run on the same indexing nodes as vector indexes. Both receive changes asynchronously through CDC. Consequently, a successful row update and that update becoming searchable are separate moments. For the support application, I would measure how long a newly added answer takes to appear in search alongside query latency.
You can put text and vector indexes on the same table, but hybrid search currently requires two queries. The application merges the result lists, for example using reciprocal rank fusion, which combines their rankings without requiring comparable raw scores.
For multi-tenant applications, another restriction matters more: a full-text query cannot include an additional WHERE restriction, including a tenant or partition key. The search FAQ confirms that vector search has filtering support, while full-text search does not.
Application filtering therefore needs careful design. Retrieving the top 20 matches globally and then removing other tenants' tickets may leave no usable results. Authorization must also happen before any retrieved content reaches a user or an AI agent. For a shared ticket table, this limitation could determine whether the built-in text search is suitable at all.
For simpler search requirements, having rows, embeddings and text indexes behind CQL can remove a separate search integration. I would still test the actual queries before planning to retire an existing search engine.
CDC gives applications a queryable change log
The search service is one consumer of ScyllaDB's change stream. Your application can consume it too.
ScyllaDB CDC records changes in log tables queried through CQL. Those records are divided into streams, so a consumer follows multiple streams and tracks its progress. The log records mutations, but you can enable optional preimages and postimages to provide additional row state (at extra cost).
Native Apache Cassandra CDC exposes per-node commit-log segments for consumers to process. Consumers must manage those files, including their removal. If the configured CDC space fills, writes to CDC-enabled tables can be rejected. ScyllaDB exposes a different integration interface: clients read database tables rather than collecting commit-log files from nodes. This is one of the strongest differences between CDC implementations. Cassandra consumers have to handle copies of each mutation, with no cross-node ordering, and must deduplicate them. ScyllaDB's log table is read through CQL, so the consumer sees one logical record.
For a Kafka pipeline, the ScyllaDB CDC Source Connector reads those log tables and publishes changes through Kafka Connect. That can feed a search engine, cache or analytical projection.
For our tickets, CDC could also trigger embedding generation after the text changes. The worker must associate an embedding with the text version it processed. Otherwise, a slow job for an older revision could overwrite the embedding for a newer one. ScyllaDB's vector index will faithfully index the vector you write, including an outdated one.
Retention needs the same attention. CDC records have a configurable lifetime. If a consumer falls behind that window, it needs a rebuild or reconciliation path. A change feed is useful only if the application knows how to recover after missing part of it.
Give the backfill a lower priority
Introducing search often means processing years of existing records. The backfill reads tickets, calls an embedding model and writes vectors while the support API continues serving users.
ScyllaDB's workload prioritization lets you attach service levels to database roles. Each can receive a different share of scheduling resources. The API and backfill can use separate roles, with the API assigned a higher priority.
Example - favour API requests over the backfill. For existing roles used by the two clients:
CREATE SERVICE_LEVEL interactive WITH SHARES = 1000;
CREATE SERVICE_LEVEL bulk WITH SHARES = 100;
ATTACH SERVICE_LEVEL interactive TO api_role;
ATTACH SERVICE_LEVEL bulk TO backfill_role;Shares assign relative priority under contention. They can favour interactive requests over backfill work, but do not reserve hardware or guarantee a particular p99.
This applies to the database operations made by those roles. A lower-priority backfill still needs its own concurrency limits, and its database service level should not be assumed to control the separate vector indexing service. I would watch API latency, backfill progress and indexing lag together when adjusting its concurrency.
Compaction changes how much storage you need
The search features are more visible, but compaction can have a larger effect on the infrastructure bill.
An LSM database periodically merges SSTables. During a large merge, input files and newly written output can coexist, so the cluster needs space beyond the live dataset.
ScyllaDB's Incremental Compaction Strategy, or ICS, divides SSTable runs into smaller fragments, 1 GB by default. Compaction processes those fragments and releases consumed input incrementally. It avoids retaining the full set of large input files until the entire operation finishes.
ICS retains the general read/write trade-offs of size-tiered compaction. Frequently updated rows can still be spread across several SSTables, and obsolete versions can occupy space until they are merged. Reducing temporary compaction space does not eliminate every form of storage overhead.
The relevant Cassandra comparison is now Unified Compaction Strategy, introduced in Cassandra 5 and covered in the earlier article. UCS supports configurable tiered and leveled behaviour with sharded compaction. A comparison based only on Cassandra's older size-tiered defaults would miss that work.
Disk usage changes as updates accumulate and compaction runs. The size of a freshly loaded dataset won’t tell you too much about the capacity required after months of overwrites, deletions and maintenance.
SSTables can also live in object storage
ScyllaDB 2026.3 introduces preview support for storing SSTables in S3-compatible object storage. A keyspace's STORAGE option selects the configured endpoint and bucket. These objects are the database's live SSTables, while compaction and reads operate on them.
This connects to the storage question from my diskless Kafka article: how much data must a compute node physically own? ScyllaDB's implementation allows tablet ownership to move by transferring references to shared objects, avoiding a copy of the SSTable payload during migration.
The current object-storage documentation targets archival workloads. Reads have substantially higher latency than local storage, and there is no local disk cache. How many SSTables a query touches therefore affects remote access costs as well as latency.
The preview requires tablets and adjustments to paging, timeouts and tablet balancing. It currently lacks secondary indexes, materialized views, CDC, LWT, counters and backup/restore support. Clusters mixing local and object-backed user keyspaces are also unsupported. The search and CDC features described earlier therefore cannot be combined with this storage mode.
Object storage also provides no automatic backup of those live SSTables: ScyllaDB deletes objects when their last reference disappears, and ScyllaDB Manager cannot back such keyspaces up. With backup missing, this is a feature to evaluate on disposable archival datasets. The ability to move tablet ownership without copying its data is worth following as the implementation develops.
A familiar interface, different internals
CQL and the wide-column data model still connect ScyllaDB to Cassandra. Underneath, the two databases make different choices about where indexes live, how applications consume changes and how SSTables are organized. Those choices affect query capabilities, data freshness and storage costs.
The previous article covered ScyllaDB's execution model and cluster management. Search, CDC and storage extend that divergence into the application architecture. Describing ScyllaDB as Cassandra in C++ leaves out an increasing share of the system.
Reviewed by: Krzysztof Ciesielski