How to Run Hybrid Search in Apex with Data 360 and ConnectApi?
Why Run Hybrid Search in Salesforce Apex?
You can run hybrid search in Salesforce Apex by sending a hybrid_search() SQL query to Data Cloud (Data 360) through ConnectApi.CdpQuery. It combines keyword precision with vector (semantic) understanding, so a query like “Expedite Order” surfaces the most relevant Knowledge article chunks even when the exact wording differs.
For consultants, admins, developers, and architects, this matters because Apex is where much custom logic lives: service consoles, Lightning Web Components, Flows that call Apex actions, and grounding logic for AI agents. In this guide you will learn how hybrid search works, see a complete Apex service class, understand each line, and finish with a list of recommendations to harden it for production.
What Is Hybrid Search in Data Cloud?
Hybrid search in Data Cloud queries a keyword index and a vector index, merges the results, and reranks them into one score. Salesforce documents the reason: vector search matches meaning but can miss specific terms and numbers, such as telling “LaserPrinter TX 400” from “LaserPrinter TX 440”, while keyword search handles those exactly (Hybrid Search).
Before a hybrid index exists, unstructured data is chunked into smaller passages. Each chunk is vectorized and keyword-indexed. That is why the code in this post touches two objects: an index table that returns ranked hits, and a chunk table that holds the readable text.
The documented syntax of the function is:
select * from hybrid_search(
table(<Search_Index_DMO>),
'<Search String>',
'<PreFilteringColumn><Operator><Value>',
<Limit Results>
);
The result columns, per the Query API documentation, are:
| Column | Meaning |
|---|---|
| RecordId__c | Unique identifier for the result; same value as SourceRecordId__c |
| SourceRecordId__c | Identifier of the source Chunk DMO record; use it to join to the chunk table |
| vector_score__c | Similarity score from the vector index |
| keyword_score__c | Similarity score from the keyword index |
| hybrid_score__c | Combined score; also reflects ranking factors such as popularity and recency |
Prerequisites and Architecture
The Apex class needs three things in place before it returns results:
- A hybrid search index in Data Cloud. Here it is built on Salesforce Knowledge articles, which produces an index DMO and a chunk DMO. Your object names will differ from the ones in the sample.
- Access to query Data Cloud from Apex. The running user must be able to query Data 360 objects; confirm the exact permissions in your org.
- The Language field and other filterable fields you plan to pre-filter on must exist on the index.
The request path is short: your Apex method builds one ANSI SQL string, hands it to ConnectApi.CdpQuery, Data Cloud runs hybrid_search() on the index DMO, and the join pulls the matching chunk text back from the chunk DMO.
[embed: node/0ae30320-1ae2]
Read left to right: Apex sends the SQL, Data Cloud ranks hits on the index DMO, and the join attaches chunk text before the rows return.
The Complete Apex Class
Here is the full KnowledgeSearchService class, exactly as written:
/**
* Service class for running hybrid (keyword + vector) searches
* against Data Cloud Knowledge Article index and chunk tables.
*/
public class KnowledgeSearchService {
/**
* Performs a hybrid search on the knowledge article index and joins the
* matching records to their chunk text, ordered by hybrid score (highest first).
*
* @param searchString The user's search text (single quotes are escaped).
* @param filterClause The filter expression passed to hybrid_search.
* @param resultLimit The maximum number of results to return.
* @return List of result rows (SourceRecordId__c, Chunk__c, hybrid_score__c).
*/
public static List<Object> performHybridSearch(String searchString, String filterClause, Integer resultLimit) {
List<Object> resultsList = new List<Object>();
// Data Cloud DMO names for the knowledge article index and chunk tables
String indexTable = 'KA_All_Knowledge_Articles_175_index__dlm';
String chunkTable = 'KA_All_Knowledge_Articles_175_chunk__dlm';
// Build the ANSI SQL query: run the hybrid search on the index table,
// then join to the chunk table to retrieve the matching chunk text
String sqlQuery = 'SELECT hs.SourceRecordId__c, ch.Chunk__c, hs.hybrid_score__c ' +
'FROM hybrid_search(TABLE(' + indexTable + '), ' +
'\'' + String.escapeSingleQuotes(searchString) + '\', ' +
'\'' + filterClause + '\', ' + resultLimit + ') hs ' +
'INNER JOIN ' + chunkTable + ' ch ON ch.RecordId__c = hs.SourceRecordId__c ' +
'ORDER BY hs.hybrid_score__c DESC';
try {
// Wrap the SQL in a Data Cloud query input
ConnectApi.CdpQueryInput queryInput = new ConnectApi.CdpQueryInput();
queryInput.sql = sqlQuery;
// Execute the query against Data Cloud
ConnectApi.CdpQueryOutputV2 response = ConnectApi.CdpQuery.queryANSISqlV2(queryInput);
// Only populate results when the response contains data
if (response != null && response.data != null) {
System.debug(response.data);
resultsList = response.data;
}
} catch (Exception ex) {
System.debug('Error executing Search Join: ' + ex.getMessage());
// Rethrow so the caller can handle the failure
throw ex;
}
return resultsList;
}
}
Code Walkthrough: How the Query Works
The method takes a search string, a filter expression, and a limit, and returns the result rows ordered by relevance.
- Declare the tables.
indexTableandchunkTablehold the API names of the Knowledge article index and chunk DMOs. Both end in__dlm, the suffix Data Cloud uses for data lake objects. - Build the SQL. The query calls
hybrid_search(TABLE(indexTable), searchString, filterClause, resultLimit)and aliases it ashs. The search text and the filter are each wrapped in single quotes, becausehybrid_search()expects them as string arguments. - Join to the chunk table.
hs.SourceRecordId__cis matched toch.RecordId__c. This follows the documented pattern: the index returns ranked hits, and the chunk DMO supplies the passage text inChunk__c. - Sort by relevance.
ORDER BY hs.hybrid_score__c DESCputs the best match first. - Wrap and run. A
ConnectApi.CdpQueryInputcarries the SQL, andConnectApi.CdpQuery.queryANSISqlV2()executes it against Data Cloud. Per the Apex reference, this method returns aConnectApi.CdpQueryOutputV2. - Guard the response. Results are copied only when
responseandresponse.dataare not null, so the method returns an empty list instead of failing on a null. - Log and rethrow. Any exception is written to the debug log and thrown again, so the calling code decides how to handle it.
One detail trips people up: the filter’s quotes. Data Cloud reads the filter as a SQL string, so an inner quote is doubled. That is why the example call passes Language__c = ''en_US'' (see the next section).
Executing the Hybrid Search from Anonymous Apex
To test the class, open the Developer Console, choose Debug, then Open Execute Anonymous Window, tick Open Log, and run:
KnowledgeSearchService.performHybridSearch(
'Expedite Order',
'Language__c = \'\'en_US\'\'',
5
);
The three arguments do the following:
'Expedite Order'is the natural-language search text. Hybrid search matches it by keyword and by meaning.'Language__c = \'\'en_US\'\''is the pre-filter. In Apex,\'is a literal single quote, so the string that reaches SQL isLanguage__c = ''en_US''. The class then wraps it in outer quotes, and the doubled inner quotes are read as one quote character. This mirrors the documented pre-filter example,RecordId__c=''b7e6...''.5limits the search to the top five hits.
In the debug log, filter on USER_DEBUG to see the payload printed by System.debug(response.data). Each result row carries the three selected columns: the source record ID, the chunk text, and the hybrid score. Your scores and text will depend on your Knowledge content, so run it against your own index and confirm that the highest hybrid score comes first.
Recommendations and Best Practices for Production
The class works as a clean proof of concept. Per your request, the code above is unchanged; these are the improvements to consider before it goes live. Items marked “verify” depend on your org and should be tested there.
| # | Area | Observation | Recommendation |
|---|---|---|---|
| 1 | Filter injection | filterClause is concatenated into the SQL without escaping. | Never pass user input into it. Build the filter from an allow-listed set of fields and values. |
| 2 | Quote escaping | String.escapeSingleQuotes() adds a backslash, while SQL string literals double the quote (''). | Test an apostrophe such as “what’s” against your index (verify). Consider replace('\'', '\'\'') if the backslash form fails. |
| 3 | Null or invalid limit | A null resultLimit concatenates as the text null and produces invalid SQL. | Validate it, apply a default, and cap it at a maximum your use case needs. |
| 4 | Hardcoded DMO names | The index and chunk names, including the numeric suffix, are specific to one org. | Move them to Custom Metadata or Custom Labels so sandboxes and production can differ without a code change. |
| 5 | Debug logging | System.debug(response.data) writes full Knowledge text to the log. | Remove it or guard it in production. Logs add volume and may expose sensitive content. |
| 6 | Untyped return | List<Object> pushes parsing onto every caller. | Return a small typed wrapper with source ID, chunk text, and score. |
| 7 | Sharing declaration | The class declares no sharing keyword. | State with sharing or inherited sharing explicitly to document intent. |
| 8 | Error context | The catch block logs and rethrows the raw exception. | Wrap it in a custom exception that carries context, and avoid logging raw search text if it can contain personal data. |
| 9 | Batches | Only the first response is read. | For larger result sets, use nextBatchId with nextBatchAnsiSqlV2 as the Apex reference describes. |
| 10 | Data space | The call uses the default data space. | If your index lives elsewhere, use the queryAnsiSqlV2(input, dataspace) overload. |
| 11 | Testability | Calling ConnectApi directly makes unit tests hard to isolate. | Put the call behind a small wrapper you can stub in tests (verify test-context behavior in your org). |
| 12 | Call volume | Each call is a Data Cloud query. | Avoid calling it per record inside loops or triggers. Batch or cache where the use case allows. |
Frequently Asked Questions
What is hybrid search in Salesforce?
Hybrid search in Data Cloud (Data 360) queries a keyword index and a vector index, merges the results, and reranks them with a combined hybrid_score__c.
How do I call hybrid_search from Apex?
Build an ANSI SQL string that selects from hybrid_search(TABLE(<index DMO>), '<text>', '<filter>', <limit>), put it in a ConnectApi.CdpQueryInput, and run it with ConnectApi.CdpQuery.queryAnsiSqlV2().
Why does the query join to a chunk table?
The index returns ranked hits with a SourceRecordId__c. The chunk DMO holds the passage text, so the join returns readable content next to each score.
How is hybrid_score__c calculated?
It combines the keyword and vector scores and can reflect ranking factors such as popularity and recency, as described in Hybrid Search Fusion Ranking.
Is Data Cloud the same as Data 360?
Yes. Salesforce documentation now uses the name Data 360 for the product also known as Data Cloud.
Conclusion
Running hybrid search in Salesforce Apex takes one SQL string, one ConnectApi.CdpQuery call, and a join from the index hits to the chunk text. Start with the class as shown, run the example against your own index, then work through the recommendations table before you ship.
Try it in a sandbox this week, and tell me in the comments which Knowledge scenarios you ground with hybrid search. Consider pairing this post with your guides on search indexes and Agentforce grounding.
Sources
- Hybrid Search, Salesforce Help
- Run Hybrid Search Queries with Query API, Salesforce Help
- Hybrid Search Fusion Ranking, Salesforce Help
- CdpQuery Class, Apex Reference
- ConnectApi.CdpQueryOutputV2, Apex Reference