[Preview] Full-text inverted index
Since version 3.3.0, StarRocks supports full-text inverted indexes, which can break the text into smaller words, and create an index entry for each word that can show the mapping relationship between the word and its corresponding row number in the data file. For full-text searches, StarRocks queries the inverted index based on the search keywords, quickly locating the data rows that match the keywords.
The full-text inverted index is not yet supported in Pramiary Key tables and shared-data clusters.
Overviewβ
StarRocks stores its underlying data in the data files organized by columns. Each data file contains the full-text inverted index based on the indexed columns. The values in the indexed columns are tokenized into individual words. Each word after tokenization is treated as an index entry, mapping to the row number where the word appears. Currently supported tokenization methods for English tokenization, Chinese tokenization, multilingual tokenization, and no tokenization.
For example, if a data row contains "hello world" and its row number is 123, the full-text inverted index builds index entries based on this tokenization result and row number: hello->123, world->123.
During full-text searches, StarRocks can locate index entries containing the search keywords using full-text inverted indexes, and then quickly find the row numbers where the keywords appear, significantly reducing the number of data rows that need to be scanned.
Basic operationβ
Create full-text inverted indexβ
Before creating a fulltext inverted index, you need to enable FE configuration item enable_experimental_gin.
ADMIN SET FRONTEND CONFIG ("enable_experimental_gin" = "true");
Also, a fulltext inverted index can only be created in the Duplicate Key table and the table property replicated_storage needs to be false.
Create full-text Inverted Index at table creationβ
Creating a full-text inverted index on column v with English tokenization.
CREATE TABLE `t` (
`k` BIGINT NOT NULL COMMENT "",
`v` STRING COMMENT "",
INDEX idx (v) USING GIN("parser" = "english")
) ENGINE=OLAP
DUPLICATE KEY(`k`)
DISTRIBUTED BY HASH(`k`) BUCKETS 1
PROPERTIES (
"replicated_storage" = "false"
);
- The
parserparameter specifies the tokenization method. Supported values and descriptions are as follows:none(default): no tokenization. The entire row of data in the indexed column is treated as a single index item when the full-text inverted index is constructed.english: English tokenization. This tokenization method typically tokenizing at any non-alphabetic character. Also, uppercase English letters are converted to lowercase. Therefore, keywords in the query conditions need to be lowercase English rather than uppercase English to leverage the full-text inverted index to locate data rows.chinese: Chinese tokenization. This tokenization method uses the CJK Analyzer in CLucene for tokenization.standard: Multilingual tokenization. This tokenization method provides grammar based tokenization (based on the Unicode Text Segmentation algorithm) and works well for most languages and cases of mixed languages, such as Chinese and English. For example, this tokenization method can distinguishes between Chinese and English when these two languages coexist. After tokenizing English, it converts uppercase English letters to lowercase. Therefore, keywords in the query conditions need to be lowercase English rather than uppercase English to leverage the full-text inverted index to locate data rows.
- The data type of the indexed column must be CHAR, VARCHAR, or STRING.
Add full-text inverted index after table creationβ
After table creation, you can add a full-text inverted index using ALTER TABLE ADD INDEX or CREATE INDEX.
ALTER TABLE t ADD INDEX idx (v) USING GIN('parser' = 'english');
CREATE INDEX idx ON t (v) USING GIN('parser' = 'english');