StarRocks version 4.0
Downgrade Notes
-
After upgrading StarRocks to v4.0, DO NOT downgrade it directly to v3.5.0 & v3.5.1, otherwise it will cause metadata incompatibility and FE crash. You must downgrade the cluster to v3.5.2 or later to prevent these issues.
-
Before downgrading clusters from v4.0.2 to v4.0.1, v4.0.0, and v3.5.2~v3.5.10, execute the following statement:
SET GLOBAL enable_rewrite_simple_agg_to_meta_scan=false;After upgrading the cluster back to v4.0.2 and later, execute the following statement:
SET GLOBAL enable_rewrite_simple_agg_to_meta_scan=true;
4.0.13β
Release Date: July 16, 2026
Behavior Changesβ
- The escape handling of
LIKEpredicates with constant operands (folded on the FE) now matches MySQL 8:SELECT 'a\\b' LIKE 'a\\\\b'returns1andSELECT 'a\\b' LIKE 'a\\b'returns0. Queries that relied on the previous non-MySQL escaping semantics will return different results. #74814 SHOW [FULL] FUNCTIONSnow always includes theisolationproperty (sharedorisolated) in the Properties column of UDFs, so users can tell whether the property is set without recreating the function. #75255- Iceberg REST catalogs with vended credentials use the table metadata cache again, reverting the earlier cache bypass that sent every
getTable()to the REST catalog and caused AWS Lake FormationRate exceededfailures. Cached tables now renew their credentials on every refresh cycle, and the table cache expiry for REST catalogs is additionally capped at 3000 seconds. #75431
Improvementsβ
- Added checksum protection for shared-data tablet metadata and transaction logs. #74924
- Supported combined transaction log / file bundling for
FRONTEND_STREAMINGloads. #74460 - Scoped shared-data schema-change job locks to the table to reduce lock contention. #75087
- Batch tablet force-delete marking now acquires the
TabletInvertedIndexwrite lock once per batch instead of once per tablet. #75616 - Added an FE metric for the maximum pending-publish time of committed transactions. #75025
- Added a memory limit check for column upgrades in window operator processing. #75821
- Removed unnecessary per-row seeks in the offsets-only read path of array columns. #75861
- Foreground row-count estimation of Iceberg tables is now computed from manifest metadata without enumerating every data file. #75280
- Addressed security vulnerabilities (CVE): excluded the vulnerable
org.jline:jline(jline-remote-telnet) from Hadoop transitive dependencies, and upgraded jackson-databind to 2.21.4. #75066 #75373
Bug Fixesβ
The following issues have been fixed:
- Incremental scan-range scheduling could recompute a different per-driver layout when reusing a deployed fragment instance, leaving part of the scan ranges unconsumed and losing rows (for example, INSERT from Hive). #74674
- INT96 timestamps nested inside ARRAY/MAP/STRUCT in Parquet files read via
FILES()or Broker Load missed the session-timezone conversion and were returned shifted by the timezone offset (top-level INT96 columns were correct). #74868 - The audit log did not record the exported row count of
SELECT INTO OUTFILE. #74467 - A strict cast could raise an overflow error from the underlying data of NULL rows, which should be ignored. #74903
parse_jsondid not respect theALLOW_THROW_EXCEPTIONsetting when handling invalid input. #74976- First-load statistics collection could not be enabled per table while disabled globally: an explicitly set table property now takes precedence over the global configuration
enable_statistic_collect_on_first_load. #74794 - Partial column updates on shared-data tables could crash the BE or silently corrupt data when the tablet schema drifted from the transaction schema. #74005
- Unexpected BE process restarts. #74424
- A BE crash (SIGFPE) in Iceberg
truncate/bucketpartition transforms when the width or bucket count is zero. #74998 - A BE crash caused by a null
driver_executorinFragmentContext::set_final_status. #75030 - A race between transaction begin and autovacuum could delete a still-needed transaction log, permanently wedging the partition's publish ("Both txn_log and corresponding tablet_meta missing"). #74906
avg(DISTINCT x)was incorrectly rewritten to use a sum/count materialized view, returning wrong results. #75071- A boundary bug in TopN RANK sorting could produce incorrect results. #75045
split/split_part/str_to_mapwith an empty delimiter could read out of bounds on invalid UTF-8 input. #75068ALTER TABLE ... MODIFY COLUMN ... AFTERa nonexistent column now returns a clear error message. #75073- A BE crash (SIGFPE) in
mod()/pmod()when computing the type's minimum value modulo -1. #74980 bar()grew memory without bound on a negative or huge width (potential DoS); such inputs are now rejected with an error. #75143- The transaction-state callback was not unregistered when a multi-statement stream load task was removed. #75188
- A crash when a query was cancelled during spill partition sorting. #75140
- The query memory limit was not enforced during table function execution. #75179
- Selecting a column named
floororceilfailed at parse time with a ClassCastException. #75241 - A heap-use-after-free in
OrderedPartitionExchangerwhen the previous chunk was mutated downstream. #75279 - Three FE metadata-lock correctness races. #74968
- Load spilling could dereference a missing query context when recording spill metrics. #75236
- ADLS2
ListPathson storage accounts without hierarchical namespace caused CN crashes and vacuum failures. #75166 - BE/CN JVM metrics emitted invalid Prometheus
# TYPElines. #75240 - JIT code generation truncated LARGEINT literals to 64 bits, producing wrong results. #75137
- A combined ALTER TABLE on an external Iceberg table re-executed already-queued actions. #74036
- A nested-loop join crash caused by a build-side column nullability mismatch. #75343
- Iceberg manifest column statistics are now cached selectively to avoid excessive FE memory usage. #75395
addPhysicalPartitioncould create only one physical partition per call, making physical-partition backfill of random-distribution tables extremely slow. #75430- A SQL injection vulnerability in the
information_schema.task_runspredicate lookup. #75520 - A BE crash (SIGSEGV) when multiple UNNEST calls in one query referenced the same array column. #75012
- Incorrect translation of nested dictionary-encoded expressions across exchange nodes. #75246
- The join skew hint was lost when the optimizer duplicated expressions. #68964
- Collecting the tables referenced by a view did not skip CTE references. #74813
- A BE abort on
CAST(json AS STRUCT<...>)when a struct field name is not a valid JSON path. #75355 - Temporary Parquet dictionary-code columns could leak to upper layers of the execution plan. #74452
- Incorrect results caused by the sort-column elimination optimization. #74983
- The UNNEST result struct type was not narrowed consistently with its input array element type. #75445
- A race between the FE EOS-cancel and the BE stage-2 deployment could mark a successfully finished query as cancelled. #75009
- Join reorder pruning could drop columns that predicates still referenced. #74791
- The global metadata lock was not released when a metadata image dump failed to acquire some database locks, blocking subsequent metadata operations. #75488
StringSearch::_patternwas left uninitialized, and thesplit_debug_symbollog output was incorrect. #75614LIKEwith the single-character wildcard_returned wrong results on columns with a GIN inverted index. #75551- The segment-iterator vector could get positionally misaligned during shared-data primary key index rebuild. #74887
- With file bundling enabled, vacuum could delete a bundle file still referenced by sibling tablets because orphaned bundled segments were not flagged as shared; subsequent publishes then failed with "Object ... does not exist". #75689
histogram()crashed on a non-positive bucket count; it now returns an error. #75041- A compaction crash (CHECK failure in
JsonMergeIterator) when a flat-JSON column changed from NOT NULL to nullable. #75680 - Common predicate operators were lost when the join tuning guide rebuilt a join. #75773
- An
UnsupportedOperationExceptioninApplyTuningGuideRulecaused by an immutable inputs list inOptExpression. #70785 SHOW PARTITIONSandpartitions_metareported the logical-partition bucket count instead of the per-physical-partition count for shared-data tables. #75734- A memory allocation exception in
NLJoinProbeOperator::pull_chunkcrashed the BE instead of failing the query. #75788 SHOW CREATE ROUTINE LOADemitted a spurious comma before the first load-property clause, making the output non-executable. #75522- CTAS (
CREATE TABLE AS SELECT) did not accept anENGINEclause, making CTAS into a Unified catalog impossible. #75771 - OR predicates over null-safe-equal (
<=>) join conditions were incorrectly rewritten toUNION ALL, returning wrong results. #75038 - A query cache normalization crash for tables with sub-partitions. #75789
array_map/transformdropped NULL rows when all non-null input arrays were empty. #75141regexp_extract_allfell into an infinite loop on a zero-length capture group. #75798- Flat-JSON subfield reads returned NULL for keys whose name contains
., which could also cause silent row loss when predicates were pushed down. #75583 - UNNEST output struct subfields were pruned by the output's own access group instead of the input array's, which could prune subfields that were still needed. #76002
- The FE Iceberg manifest data-file cache could serve an incomplete file set, silently dropping a data file from scan planning and under-counting query results. #76215
- The Iceberg table metadata cache was not invalidated after an INSERT commit, so subsequent queries could miss the latest snapshot. #67230
4.0.12β
Release Date: June 25, 2026
Behavior Changesβ
- When reading INT64 timestamps from Parquet files written with
isAdjustedToUTC=false(timezone-naive),SELECT FROM FILES()and broker/stream LOAD no longer shift the values by the session timezone offset. Such timestamps are now read as wall-clock values, consistent with Trino, Spark, and Impala. Previously the values drifted whenever the session timezone was not UTC. #73674 - CTAS (
CREATE TABLE AS SELECT) now preserves the declaredVARCHAR(N)length when the source carries an explicit user length (a catalog column reference,CAST AS VARCHAR(N), or a string literal), instead of widening it toVARCHAR(1048576). This keeps the length constraint enforceable and aligns DDL with dbt schema contracts. Materialized view materialization still widens columns as before. #73498 - The Paimon connector now respects the session variable
connector_max_split_sizewhen calculating scan splits, instead of always using the default value, so tuning it now affects Paimon scan parallelism. #71756
Improvementsβ
- Optimized
base64_to_bitmapby folding the conversion at constant-evaluation time for constant inputs. #74684 ngram_searchnow supports a non-constant needle (the search term can be a column expression rather than only a constant). #74675- The Arrow-to-JSON converter now supports
LARGE_LISTandFIXED_SIZE_LISTtypes. #73714 - Added an opt-in option to isolate wide-string columns during statistics collection to reduce memory pressure. #73258
information_schema.COLUMNSnow populates theDATETIME_PRECISIONfield. #74623- Relaxed database read locks to table-scoped intensive locks in
InformationSchemaDataSourceandFrontendServiceImplto improve concurrency. #73936 #73913 - Narrowed database write locks to table-scoped intensive write locks for shared-nothing clusters, and scoped replica row-count updates to the table lock. #74523 #74521
- Moved the routine-load broker RPC out of the per-job write lock to reduce contention. #73591
- Deferred JDBC
REMARKSfetching out of thegetTable()hot path to speed up metadata access for JDBC catalogs. #73488 - Pushed down the
table_namepredicate forinformation_schema.tables_configqueries. #73210 - Skipped per-replica scans on single-medium BEs in
BackendLoadStatistic. #73555 - Added a write timeout to the MySQL channel result send path to prevent stuck connections. #73646
- Added catalog recycle bin size gauge metrics. #74440
- Added vacuum batch-size and retry-count metrics, and added decorrelated jitter to the lake vacuum retry backoff to reduce retry storms. #74112 #74108
- Upgraded third-party dependencies to address security vulnerabilities (CVE): Netty to 4.1.135.Final, Tomcat to 9.0.118, and Thrift to 0.23.0. #74668 #73797 #73625
Bug Fixesβ
The following issues have been fixed:
- Successfully committed multi-statement transaction stream loads were shown as
PREPARINGforever ininformation_schema.loadsandSHOW STREAM LOAD. #74386 - Rows were silently dropped from
information_schema.loadson clusters whose session timezone differs from Asia/Shanghai, because load times were exchanged as naive wall-clock strings across the BE/FE thrift boundary. #73365 - The
COMMITof an explicit transaction waited onlyquery_timeoutmilliseconds (instead of seconds) for the database write lock due to a unit mismatch. #73549 current_timestamp/now()column defaults were displayed as a frozen literal afterALTER TABLE ... ADD COLUMNand could be lost across FE restarts or edit-log replay. #73455- Querying
sys.fe_memory_usage/sys.fe_lockswithout theOPERATE ON SYSTEMprivilege returned a misleading RPC-failure message instead of a clear access-denied error. #73567 - Automatic per-key Hive partition stats refresh could overload the Hive Metastore for tables with many partitions. #73563
- A null-pointer issue when reading the GTID during a schema change. #74855
- An empty analytic operator was not pruned after pushing down a distinct aggregation. #74810
- Zero row counts could corrupt partition statistics. #74801
- Vector index rewrite could pollute the shared table schema. #74785
- An
IllegalStateExceptionduring parallel profile collection, fixed by making Tracers fork-aware. #74746 - BE vacuum tasks were not aborted once the FE caller's timeout elapsed. #74694
- Partition consumer errors in
ChunksPartitionerwere lost instead of being propagated. #74693 - A lock mismatch in
blockingAddTabletCtxToScheduler. #74596 - A typo in the
azure_adls2_oauth2_client_endpointconfiguration field name. #74581 - Pipeline observers were not notified on missed operator state transitions. #74557
- The reported vacuum watermark was incorrect when retain-boundary metadata was gone. #74429
- A data race on
MaterializedIndexMetaduringupdateSchemaBackendId. #74412 - A non-primary-key replica could get stuck with a permanent version hole; it now self-heals. #74408
- A use-after-free of
LLVMContextwhen JIT compilation fails. #74396 - A column mismatch in the missing-replica row of
ADMIN SHOW REPLICA STATUS. #74393 - Invalid JIT IR generated for
CASE WHENwith mixed float/int WHEN and result types. #74382 - The
CatalogRecycleBinwas frozen when a cluster snapshot kept failing. #74379 - A partial update targeting a table modified earlier in the same explicit transaction is now rejected with a clear error. #74344
- Immutable-partition updates did not use the transaction's compute resource. #74316
- A potential out-of-bounds error caused by partitioned join. #74315
- Database-level UDFs were not restored to the renamed target database for FE followers. #74313
- Non-root compound predicates yielded
EOFinstead ofNotPushDown. #74218 - Table names were not backquoted when persisting the routine load
origStmt. #74188 - An assertion name lookup error in assert-num-rows. #74178
- Aggregation used type-mismatched aggregate functions. #74159
- Force-killed task runs were not archived, and session-prefixed task-run timeouts were not honored. #74146
RENAMEandSWAP(for tables and materialized views) now take the database write lock to avoid concurrent-modification issues. #74100- Composite-rowset stats were not summed when batching op_writes in primary-key multi-statement transactions. #74059
- The sink was not notified when a distinct aggregate source finished. #74055
pipe_file_listwas not recreated when_statistics_was dropped. #73970- A crash in
TabletInvertedIndex.deleteTablets, fixed by fast-path skipping empty input. #73955 - The task manager could write an illegal edit log for a task run. #73882
- A
set_thread_namerace on data directory load threads. #73862 - A race condition in
TabletSinkSender::_send_chunk_by_node. #73820 - Incorrect memory accounting in
OlapTableSink. #73807 - Incorrect connector bytes-read statistics. #73799
BACKUP ON (ALL FUNCTION)/(ALL EXTERNAL CATALOGS)failures. #73790NullableColumnUnaryFunctionlost the decimal scale when all values were null. #73789- A FlatJSON crash when subwriters had no appends. #73730
LargeList/FixedSizeListcould not be converted to a JSON column during broker load. #73718- Partial-append of nested types failed during JSON load. #73715
- An NPE in
StatisticsCalcUtilswhen a partition is dropped concurrently. #73711 - Decimal-valued unit counters were not parsed in
RuntimeProfileParser. #73683 - An NPE during nested materialized view refresh. #73644
- Multi-statement transactions were not handled correctly in lake
publish_log_version. #73423 - An empty
ALTER TABLEclause is now rejected, andOPTIMIZEreplay was improved. #73352 - Disk cache overflow when Iceberg metadata entries are pinned. #71651
4.0.11β
Release Date: June 5, 2026
Behavior Changesβ
get_json_stringand the otherget_json_*functions now return the JSON parse error instead of NULL when implicit VARCHAR-to-JSON parsing fails underALLOW_THROW_EXCEPTION. The default behavior (returning NULL when the mode is disabled) is unchanged. #73199pipeline_enable_large_column_checkeris now enabled by default. #72798
Improvementsβ
- Lake write-path load spill files now use a flat, single-level directory layout with the transaction ID baked into each filename, and are reclaimed by a txn-id-based vacuum pass. This moves bulk deletes off the write hot path and lets vacuum clean up spill files leaked by BE crashes. #73064
- SHOW statements (such as
SHOW GRANTSandSHOW WAREHOUSES) are now allowed inside an explicit transaction, so BI/JDBC clients that automatically issue SHOW no longer break the transaction flow. #72954 - Java UDAF and UDTF now support STRUCT arguments and return types. #72911
- Scalar Java UDF now supports STRUCT arguments. #72620
- Java UDF now supports DATE and DATETIME types. #72337
- Java UDF now supports nested ARRAY/MAP types. #72283
- Added the FE configuration
deploy_serialization_min_thread_pool_size. #72274 - Skipped redundant partition key expression building when an
add_partition_valuededuplication hit occurs. #73156 - Avoided a redundant
latestSnapshot()call inPaimonMetadata#getTableVersionRange. #72892 - Deduplicated commutative AND/OR expressions in scalar operator common subexpression elimination. #72823
Bug Fixesβ
The following issues have been fixed:
- A memory leak introduced by the UDAF cache. #74025
- An incorrect implementation in aggregate combined functions. #74169
- An issue in shared-data combined txn log mode where the per-partition coordinator claim was not re-recorded on every sender's open, which could drop txn logs. #73962
- A read failure on Iceberg tables that use a custom
LocationProvider, fixed by lazily initializing theLocationProviderinSerializableTable. #73482 - A serialization failure caused by the
de.javakaffeeUnmodifiableCollectionsSerializer, now replaced with a Java 17-compatible version. #73458 HdfsFsManagercopy error messages now include the underlying cause. #73414- A concurrent
SegmentFlushTaskrace inDeltaWriter::commit(). #73371 - Sort merge provider errors are now propagated to the fragment context instead of being lost. #73337
- An issue where Ranger row-filter/masking policies on Hive views were skipped, so policies on the view or its base tables were not applied. #73265
- Upgraded libthrift to 0.23.0 to address a security vulnerability (CVE). #73243
- An FE file-descriptor leak, fixed by reusing
HttpClientinstances. #73239 - Parquet broker load errors now include file/column/row context. #73236
- A slot lookup failure for output slots with an empty
col_namein the Spark connector external scan. #73225 - A crash in
SinkBufferduring graceful exit. #73202 - Query cache conflicts with local shuffle aggregation. #73194
- A use-after-free of the Hive partition descriptor across fragment teardown. #73176
- A thread-safety issue in lake vacuum, fixed by using
localtime_r. #73088 - A race condition between
PipelineTimerTaskdoRunand unscheduling during query context destruction. #73082 - Lock contention on read-only query-engine paths, reduced by relaxing DB locks. #73067
- An materialized view refresh failure with SQL Server tables in a JDBC catalog. #72962
- A JNI local-reference leak in
JDBCScanner::_init_jdbc_scanner. #72913 - An issue where partition TopN could lose a child's output column. #72848
- An incorrect plan caused by not clearing
LambdaArgument.transformedOpbefore INSERT OVERWRITE re-planning. #72832 - The coordinator lock was held during external resource cleanup. #72830
Lockerrollback is now exception-safe and the unlock order is fixed. #72789- An incorrect byte order in
ColumnDict.merge, now using unsigned byte order. #72778 - A stack-buffer-overflow when formatting into a temporary
std::string. #72728 - The HAVING clause is now checked when disabling aggregation spill on a small LIMIT. #72705
- A hang caused by joining forwarded RPCs when draining the runtime_filter worker. #72626
- Incorrect lazy-materialization slot nullability for a materialized view over an outer join. #72621
merge_conditionwas not preserved when applying a normal rowset commit. #72542- Lock contention in
TabletScheduler/TabletSchedCtxhot paths during clone, reduced by relaxing DB locks. #72475 Lockerdid not roll back a partial intensive-lock acquisition. #72423- A spillable hash join probe crash. #72397
- COALESCE children are now cast to a common type in the JOIN USING transformer. #72338
- DB READ lock was held too broadly for single-table proc directories, now relaxed to per-table. #72334
- A memory leak when caching the materialized view plan context. #72300
- FSE-v2 did not set the schema for shared-data sorted schema change. #72235
ConsistencyCheckerheld a DB READ lock too broadly in periodic scans, now relaxed to per-table READ. #72218- A BE crash when querying
information_schema.warehouse_queries. #72019 - A trailing
\rwas not stripped before the closing enclose in CRLF CSV inputs. #71866 - Paimon primary key columns were incorrectly marked as non-nullable when querying an external catalog. #71660
- A redundant double slash was created when constructing the JDBC URL if the URI already ended with a trailing slash, breaking strict drivers such as ClickHouse. #70992
4.0.10β
Release Date: May 9, 2026
Behavior Changesβ
- Cloud storage credentials are now redacted in error messages produced by
INSERT INTO FILES, preventing accidental exposure of secrets in error logs andSHOW LOADoutput. #71245 - StarRocks no longer permits queries against insert-only ACID Hive tables in Hive catalog. Previously such queries could silently return more rows than actually visible because INSERT OVERWRITE operations were not recognized. Affected tables now return an explicit error instead of incorrect results. #71460
Improvementsβ
- Added an Avro schema cache in Iceberg
PartitionDataconstruction to remove redundant JacksonObjectMapperallocations during partition load on tables with many partitions. #72215 - Optimized
CatalogRecycleBin.getAdjustedRecycleTimestampto avoid rebuilding the table-id map on every call, reducing recycle-bin cleanup and tablet scheduling overhead. #72128 OlapTableSink.createLocationnow batches tablet-location lookups in shared-data mode, removing per-tablet StarOS RPCs that previously stalled the planner critical section. #72041- Java UDAF instances are now loaded and initialized once per query and reused across pipeline driver instances, removing the linear driver-preparation overhead at high
pipeline_dop. #72038 - Added BE metrics
starrocks_be_staros_shard_info_fallback_totalandstarrocks_be_staros_shard_info_fallback_failed_totalto track when the StarOS worker falls back to fetching shard info fromstarmgrbecause the local cache missed. #71620 - File-bundle writes now prefer a tablet-local aggregator so the bundled tablet metadata path does not require cross-node shard-info lookups. #71613
- Audit log entries now include the queried tables and views referenced by each query. #71596
INSERT INTO FILESCSV export now supportscsv.encloseandcsv.escapeproperties for controlling field quoting and escaping. #71589- Added LDAP direct bind authentication via DN pattern, removing the requirement for an admin search account in single-tenant LDAP setups. #71559
- Added the
starrocks_fe_tablet_nummetric for shared-data clusters to match the shared-nothing metric set. #71444 star_mgr_meta_sync_interval_secis now runtime-mutable viaADMIN SET FRONTEND CONFIG; the new interval takes effect on the next sync cycle without an FE restart. #71675
Bug Fixesβ
The following issues have been fixed:
- A race in shared-data combined txn log mode where INSERT into per-partition coordinator dispatch could classify legitimate txn logs as orphan and drop them, leaving the transaction stuck in non-VISIBLE state. #72237
- An issue where
_incremental_open_node_channelchannels in shared-data combined txn log mode silently dropped txn logs because the legacy "sender_id == 0 collects all logs" rule did not apply to incremental channels. #71992 - An issue where
RuntimeProfile::to_thrift()could crash BE withstd::bad_optional_accesswhen another thread reset counter min/max values during profile serialization. #72904 - An inconsistency in flat JSON merge results when one side contributed empty values. #72973
- An issue where
CREATE TABLEfor an Iceberg table failed with "Multiple entries with same key: format-version" when the user explicitly specifiedformat-versioninPROPERTIES. #72828 - A
CompactionScheduler.startCompactionlock scope that held a DB-wide READ lock across single-table critical work, blocking concurrent DDL on other tables in the same database. Switched to IS on DB plus READ on the target table. #72178 - An issue where
StarMgrMetaSyncer.syncTableMetaInternalandsyncTableColocationInfoheld DB READ/WRITE locks across external StarOS RPCs, freezing CREATE/DROP/ALTER/RENAME on every table in the database for the duration of each RPC. #72108 - An issue where
StarMgrMetaSyncer.getAllPartitionShardGroupIdheld the DB READ lock for full iteration over all cloud-native tables and physical partitions, stalling FE threads waiting for the DB write lock on large catalogs. #71614 - A redundant DB READ lock in
getTableNamesViewWithLock. The underlyingnameToTableis aConcurrentHashMap, so the enclosing lock added contention without correctness benefit. #72042 - A DB WRITE lock in the read-only
/api/{db}/{table}/_countREST endpoint that was unnecessary for computingproximateRowCount(). #72053 - A batch publish deadlock caused by partition version gaps that operations like tablet split, schema change, and alter jobs reserved by advancing
nextVersionwithout a matching publish. #71483 - A deadlock in shared-nothing mode when warming up the LRU cache for rowset metadata while the cache was full. #71459
- A
PipelineTimerTaskthat could remain stuck inwaitUtilFinisheddue to incorrect ordering between consumer registration and finished signaling. #72058 - A condition race in
ConnectorSinkPassthroughExchanger::acceptthat crashed BE with SIGSEGV via out-of-bounds vector access on_writer_count. #71848 - A use-after-free in
LoadChannel::get_load_replica_statuscaused by destruction of a temporaryshared_ptr. #71843 - A use-after-free in the information schema sink due to a missing reference count increment in async RPC closure handling. #71513
- A BE crash in
reverse(DecimalV3)caused by improper handling of decimal value width. #71834 - A BE crash when
UNNESTproduced columns whose define-expression carried an ARRAY type, which was incompatible with global dictionary generation downstream. #72027 - An NPE in FE when creating an Iceberg external table with invalid transform argument order such as
bucket(4, region); FE now returns a normal analyzer error. #71917 - An issue where Iceberg manifest data file cache entries were missing column statistics when the first query against a table did not request stats (for example
SELECT *). #71913 - An issue where the Iceberg min/max optimization was silently skipped when the table was partitioned by
bucket(col, N)becausePruneHDFSScanColumnRuleinjected a placeholder materialized column. #71863 - An issue where
AggregateJoinPushDownRulefailed to rewrite materialized views over Iceberg base tables becauseTable.getId()was compared instead of identity, and connector-table ids can shift across plan rebuilds. #71856 - An issue where INSERT OVERWRITE into Hive dynamic partitions failed when the metastore listed a partition whose location no longer existed on the file system; the missing partition directory is now created before commit. #71810
- A Parquet scanner failure (
Illegal converting from arrow type(dictionary) ...) when Arrow returned dictionary-typed columns, including dictionaries nested inside arrays, structs, and maps. #71855 - An issue where stale scan ranges from earlier batches persisted across
ColocatedBackendSelector.Assignmentincremental batches, causing files to be re-deployed and re-scanned. #71789 - An issue where
PruneShuffleColumnRuledid not update the JoinoutputPropertyafter pruning Exchange shuffle columns, leading to incorrect downstream distribution. #72003 - Incorrect shuffle distribution caused by a missing project node when
PushDownJoinOnExpressionToChildProjectwas disabled during the first stage of multi-stage materialized view rewrite. #71075 - Duplicate
Applyattachments inReplaceSubqueryRewriteRulewhen predicate normalization made the same scalar-subquery placeholder appear multiple times. #71155 - A short-circuit issue in
EventSchedulerwhere a finished join probe could prevent the pipeline from transitioning to the finished state. #71740 - An issue where AWS assume-role configured via
aws.s3.iam_role_arnwas not applied to JNI scanners (RCFile / Avro / SequenceFile / Hudi), causing S3 403 errors. #71422 - An issue where Oracle JDBC predicate pushdown produced invalid SQL because date literals did not match the Oracle NLS format; literals are now emitted as
date '...'. #71412 - An issue in shared-data mode where a follower FE forwarded DDL to the leader and waited only for FE journal replay, missing the StarMgr journal and producing "no queryable replica" errors for queries that immediately followed table creation. #71263
- An issue where
get_tablet_statsfor Primary Key tablets repeatedly reloaded the entireTabletMetadatafor every segment viaget_del_vec_in_meta(). #71672 - An Arrow Flight issue where empty result sets returned column names of
rbecause the placeholder name was emitted instead of the actual schema. #71534 - An issue where
parallel_clone_task_per_pathupdates did not include the store-path count when resizing the CLONE thread pool. #71484 - An issue where the resource group user classifier rejected digit-leading usernames that
CREATE USERallowed. The classifier now uses the same validation rule asCREATE USER. #71470 - An issue where
HttpServerHandler.channelInactiveskippedunregisterConnectionwhenisRegistered()was false, leaking connection-map entries for early-failing requests. #72006 - An issue where Java UDF JNI calls (
NewObject,NewArray,NewStringUTF, etc.) did not check for exceptions or null returns, leading to silent failures or undefined behavior. #71734 - An issue where
be_tablets.DATA_SIZEreportedtotal_disk_size(including rowset-embedded indexes and the persistent PK index for lake PK tablets) instead of rowset column data bytes. #70735 - A noisy "Failed to batch drop tablets" warning printed by
StarMgrMetaSyncereven when there were no shards to delete. #72209 - CVE-2026-42198 (pgjdbc) and CVE-2026-5598 (BouncyCastle): bumped
org.postgresql:postgresqlto 42.7.11 and BouncyCastle to 1.84. #72797 - CVE in netty: upgraded netty to 4.1.133.Final. #72905
- Cleaned broker CVEs by upgrading netty / jetty / awssdk / jackson dependencies in the broker. #72184
- Upgraded jetty-http to 9.4.58.v20250814 to address known CVEs in the previous jetty-http version. #71762
- Temporarily masked CVE-2026-2332 to unblock the build, since jetty 9.x is EOL and no upstream fix is published. #71914
4.0.9β
Release Date: April 16, 2026
Behavior Changesβ
- When VARBINARY columns appear inside nested types (ARRAY, MAP, or STRUCT), StarRocks now correctly encodes the values in binary format in MySQL result sets. Previously, raw bytes were emitted directly, which could break text-protocol parsing for null bytes or non-printable characters. This change may affect downstream clients or tools that process VARBINARY data inside nested types. #71346
- Routine Load jobs now automatically pause when a non-retryable error is encountered, such as a row causing the Primary Key size limit to be exceeded. Previously, the job would retry indefinitely because such errors were not recognized as non-retryable by the FE transaction status handler. #71161
SHOW CREATE TABLEandDESCstatements now display the Primary Key columns for Paimon external tables. #70535- Cloud-native tablet metadata fetch operations (such as
get_tablet_statsandget_tablet_metadatas) now use a dedicated thread pool instead of the sharedUPDATE_TABLET_META_INFOpool. This prevents metadata fetch contention from impacting repair and other tasks. The new thread pool size is configurable via a new BE parameter. #70492
Improvementsβ
- Added session variables to control the encoding behavior of VARBINARY values in MySQL protocol responses, providing fine-grained control over binary result encoding in client connections. #71415
- Added a
snapshot_meta.jsonmarker file to cluster snapshots to support integrity validation before snapshot restoration. #71209 - Added warning logs for silently swallowed exceptions in
WarehouseManagerto improve observability of silent failures. #71215 - Added metrics for Iceberg metadata table queries to support performance monitoring and diagnosis. #70825
- The
regexp_replace()function now supports constant folding during FE query planning, reducing planning overhead for queries with constant string arguments. #70804 - Added categorized metrics for Iceberg time travel queries to improve monitoring and performance analysis. #70788
- Added log output when update compaction is suspended, improving visibility into compaction lifecycle. #70538
SHOW COLUMNSnow returns column comments for PostgreSQL external tables. #70520- Added support for dumping query execution plans when a query encounters an exception, improving diagnosability of runtime failures. #70387
- Tablet deletion during DDL operations is now batched, reducing write lock contention on tablet metadata. #70052
- Added a Force Drop recovery mechanism for synchronous materialized views that are stuck in an error state and cannot be dropped through normal means. #70029
Bug Fixesβ
The following issues have been fixed:
- An issue where the profile
START_TIMEandEND_TIMEwere not displayed in the session timezone. #71429 - A shared-object mutation bug in
PushDownAggregateRewriterwhen processing CASE-WHEN/IF expressions, which could cause incorrect query results. #71309 - A use-after-free bug in
ThreadPool::do_submittriggered when thread creation fails. #71276 - An issue where
information_schema.tablesdid not properly escape special characters in equality predicates, causing incorrect results. #71273 - An issue where the materialized view scheduler continued to run after the materialized view became inactive. #71265
- Fixed a task signature collision in
UpdateTabletSchemaTaskacross concurrent ALTER jobs that could cause schema update tasks to be skipped. #71242 - An issue where row count estimation produced NaN values for histograms that contained only MCV (Most Common Values) entries. #71241
- A missing dependency on the AWS S3 Transfer Manager in the AWS SDK integration. #71230
- An issue where
TaskManagerscheduler callbacks did not verify whether the current node is the leader, potentially causing duplicate task execution on follower nodes. #71156 - A thread-local context pollution issue where
ConnectContextinformation was not cleared after a leader-forwarded request completed. #71141 - An issue where the partition predicate was missing in short-circuit point lookups, causing incorrect query results. #71124
- A NullPointerException when analyzing generated columns during Stream Load or Broker Load if a column referenced by the generated column expression was absent from the load schema. #71116
- A use-after-free bug in the error handling path of parallel segment and rowset loading. #71083
- An issue where delvec orphan entries were left behind when a write operation preceded compaction in the same publish batch. #71049
- An issue where queries appeared in the
current_queriesresult via HTTP loopback when checking query progress internally. #71032 - CVE-2026-33870 and CVE-2026-33871. #71017
- A read lock leak in
SharedDataStorageVolumeMgr. #70987 - An issue where the input and result columns of the
locate()function shared the same NullColumn reference inside BinaryColumns, causing incorrect results. #70957 - An issue where safe tablet deletion checks were incorrectly applied during ALTER operations in share-nothing mode. #70934
- A race condition in
_all_global_rf_ready_or_timeoutthat could prevent global runtime filters from being applied correctly. #70920 - An int32 overflow in the
ACCUMULATEDmetric macro that caused metric values to silently overflow. #70889 - Incorrect aggregation results in dictionary-encoded merge GROUP BY queries. #70866
- CVE-2025-54920. #70862
- A potential data loss issue in aggregation spill caused by incorrect hash table state handling during
set_finishing. #70851 - An issue where the
content-lengthheader was not reset whenproxy_pass_request_bodyis disabled. #70821 - An issue where the spill directory for load operations was cleaned up in the object destructor rather than during
DeltaWriter::close(), potentially causing premature deletion of spill data. #70778 - An issue where
INSERT INTO ... BY NAMEfromFILES()did not correctly push down the schema for partial column sets. #70774 - An issue where connector scan nodes did not reset the scan range source on query retry, causing incorrect results upon retry. #70762
- A potential rowset metadata loss for Primary Key model tablets caused by a GC race during disk re-migration of the form AβBβA. #70727
- An issue where a query-scoped warehouse hint leaked the
ComputeResourceobject inConnectContext, potentially affecting subsequent queries on the same connection. #70706 - An issue where redundant conjuncts in
MySqlScanNodeandJDBCScanNodecaused BE errors related toVectorizedInPredicatetype mismatches. #70694 - A missing
libssl-devdependency in the Ubuntu runtime environment. #70688 - An issue where Iceberg manifest cache completeness was not validated on read, leading to incorrect scan results when the cache was partially populated. #70675
- A duplicate closure reference in
_tablet_multi_get_rpcthat could cause use-after-free. #70657 - Partial manifest cache writes in the Iceberg
ManifestReaderthat could result in incomplete cache entries and incorrect scan behavior. #70652 - A crash in
array_map()when processing arrays that contain null literal elements. #70629 - A stack overflow in the
to_base64()function when processing large inputs. #70623 - An issue where
INSERT INTO ... BY NAMEfromFILES()used positional column mapping instead of name-based mapping, causing data to be written to incorrect columns. #70622 - An issue where
NOT NULLconstraints were incorrectly pushed down into the schema inferred fromFILES(), causing load failures for nullable columns. #70621 - An issue where precise external materialized view refresh did not fall back correctly for Iceberg-like connectors. #70589
- A
num_short_key_columnsmismatch when constructing a partial tablet schema, which could cause data read errors. #70586 - A BE crash that occurred when the child iterator was exhausted in
MaskMergeIterator. #70539 - An issue where materialized view refresh jobs repeatedly refreshed partitions whose corresponding Iceberg snapshots had expired. #70523
- An issue where starlet configuration parameters could not be set. #70482
- An issue where the lock-free materialized view rewrite path incorrectly fell back to live metadata, causing inconsistent rewrite behavior. #70475
- An issue in
JoinHashTable::merge_htwhere dummy rows were not skipped for expression-based join key columns, causing incorrect join results. #70465 - An incorrect equality comparison in
InformationFunctionthat could produce wrong results in certain queries. #70464 - A column type mismatch in the
__iceberg_transform_bucketinternal function. #70443 - An issue where Iceberg materialized view refresh failed when Iceberg snapshot timestamps were non-monotonic. #70382
- An issue where user authentication credentials were exposed in audit logs and SQL redaction output. #70360
- A CN crash that occurred when scanning an empty tablet with physical split enabled. #70281
- An issue where the VARCHAR column length was not preserved after a redundant CAST was eliminated during query optimization. #70269
- An issue where brpc connection retry logic did not correctly handle a wrapped
NoSuchElementException, causing connection failures after the retry attempt. #70203 - An issue where null fractions for outer join columns were not preserved during statistics estimation, leading to suboptimal query plans. #70144
- A memory tracker leak in connector sink operations running on poller threads. #70121
4.0.8β
Release Date: March 25, 2026
Behavior Changesβ
- Improved
sql_modehandling: whenDIVISION_BY_ZEROorFAIL_PARSE_DATEmode is set, division by zero and date parse failures instr_to_date/str2datenow return an error instead of being silently ignored. #70004 - When
sql_modeis set toFORBID_INVALID_DATE, invalid dates inINSERT VALUESclauses are now correctly rejected instead of being bypassed. #69803 - Expression partition generated columns are now hidden from
DESCandSHOW CREATE TABLEoutput. #69793 - Client ID is no longer included in audit logs. #69383
Improvementsβ
- Added a configuration item
local_exchange_buffer_mem_limit_per_driverto limit the local exchange buffer size todop * local_exchange_buffer_mem_limit_per_driver. #70393 - Cached file existence check results across versions in
check_missing_filesto reduce redundant storage I/O. #70364 - Allowed disabling split and reverse scan ranges for descending TopN runtime filters when
desc_hint_split_rangeis set to β€ 0. #70307 - Added
EXPLAINandEXPLAIN ANALYZEsupport forINSERTstatements in the Trino dialect. #70174 - Optimized Iceberg read performance when position deletes are present. #69717
- Optimized materialized view best-selector strategy based on distributed keys to improve materialized view selection accuracy. #69679
Bug Fixesβ
The following issues have been fixed:
- JDBC MySQL pushdown failing for unsupported cast operations. #70415
- Type mismatch issues in materialized view refresh. Added
mv_refresh_force_partition_typeconfiguration to force partition type in materialized view refresh. #70381 dataVersionnot set correctly when restoring from backup. #70373- Duplicated partition names in materialized view refresh tasks. #70354
- Incorrect SLF4J parameterized logging using string concatenation instead of placeholder arguments. #70330
- Comment not set when creating Hive tables. #70318
FileSystemExpirationCheckerblocking on slow HDFS close operations. #70311- Distribute column validation not applied across different partitions in
OlapTableSink. #70310 - Constant folding producing INF instead of an error when double addition overflows. #70309
- Typo in Iceberg table creation: field
commonwas used instead ofcomment. #70267 - Root user not bypassing all Ranger permission checks in some scenarios. #70254
query_poolmemory tracker going negative during data ingestion. #70228AuditEventProcessorthread exiting due toOutOfMemoryException. #70206SplitTopNRulenot applying partition pruning correctly. #70154- Out-of-bounds access in
cal_new_base_versionduring schema change publish. #70132 - Materialied view rewrite ignoring dropped partitions from the base table. #70130
- Unexpected partition predicate pruning due to type mismatch in boundary comparisons. #70097
str_to_datelosing microsecond precision in BE runtime. #70068- Join spill process crashing in
set_callback_function. #70030 - Broker Load failing GCS authentication after
gcs-connectorupgrade to version 3.0.13. #70012 - DCHECK failure in
DeltaWriter::close()when called from a bthread context. #69960 - Use-after-free race condition in
AsyncDeltaWriterclose/finish lifecycle. #69940 - Race condition causing write transaction edit log entry to be missed. #69899
- Known CVE vulnerabilities. #69863
- Follower FE not waiting for journal replay in
changeCatalogDb. #69834 - Incorrect
LIKEpattern matching with backslash escape sequences. #69775 - Expression analysis failure after renaming a partition column. #69771
- Use-after-free crash in
AsyncDeltaWriter::close. #69770 - Crash in local partition TopN execution. #69752
- Incorrect behavior in
PartitionColumnMinMaxRewriteRulecaused byPartition.hasStorageData. #69751 - Duplicated CSV compression suffix in file sink output filenames. #69749
lake_capture_tablet_and_rowsetsnot gated behind an experimental configuration flag. #69748- Incorrect partition min pruning with shadow partitions. #69641
- Java UDTF/UDAF crashing when method parameters use generic types. #69197
- Per-query metadata not released after query planning, causing FE OOM during concurrent query execution. #68444
- Query-scope warehouse hint leaking
ComputeResourceinConnectContext. #70706 - Lock-free materialized view rewrite incorrectly falling back to live metadata. #70475
- Duplicate closure reference in
_tablet_multi_get_rpc. #70657 - Infinite recursion in
ReplaceColumnRefRewriter. #66974 NOT NULLconstraint incorrectly pushed down toFILES()table function schema. #70621num_short_key_columnsmismatch in partial tablet schema. #70586COLUMN_UPSERT_MODEchecksum error in shared-data clusters. #65320- Column type mismatch for
__iceberg_transform_bucket. #70443 - Starlet configuration items not taking effect. #70482
- DCG data not read correctly when switching from column mode to row mode in partial update. #61529
4.0.7β
Release Date: March 12, 2026
Behavior Changeβ
- Disallowed creating materialized views based on Iceberg views. #69471
- Fixed inconsistencies in multi-statement Stream Load transaction behavior. #68542
Improvementsβ
- Added fine-grained trace counters for
LakePersistentIndexin the Publish phase. #69640 - Triggered early flush in
LakePersistentIndexwhen rebuild row count exceeds the threshold. #69698 - Added
dump_lake_persistent_index_sstoperation tometa_tool. #69682 - Improved
REPAIR TABLEfunctionality andSHOW TABLETstatus display. #69656 - Supports
ADMIN SHOW TABLET STATUSfor cloud-native tables. #69616 - Upgraded
hadoop-clientfrom 3.4.2 to 3.4.3. #69503 - Prevented crashes when deserialization mismatches occur. #69481
- Pushed down predicates to FE when querying
information_schema.loads. #69472 - Optimized the SQL displayed for materialized view refresh TaskRuns. #69437
- Bypassed caching in
CachingIcebergCatalogwhen vended credentials are enabled. #69434 - Uses
tryLockwith timeout incanTxnFinishedto reduce lock contention. #69427 - Added a global readonly variable
@@run_mode. #69247 - Uses Estimator to estimate cache entry weight in
DeltaLakeMetastore. #69244 - Added resource share type support for AWS Glue
GetDatabasesAPI. #69056 - Extracted range predicates from scalar-subqueries containing
convert_tz. #69055 - Added ByteBuffer Estimator. #69042
- Supports fast cancel for Lake DeltaWriter in shared-data clusters. #68877
- Supports an interface to add physical partitions for random distribution tables. #68503
- Gated SQL transactions behind the session variable
enable_sql_transaction(default: true). #63535 - Added partition scan number limit when querying external tables. #68480
Bug Fixesβ
The following issues have been fixed:
- Incorrect value for the metric
g_publish_version_failed_taskswhen the resource is busy. #69526 - NPE in
IcebergCatalog.getPartitionLastUpdatedTimewhen the snapshot has expired. #68925 - DCHECK failure in
DeltaWriter::close()when called from bthread context. #70057 - Several use-after-free issues. #69968
- Use-after-free race in AsyncDeltaWriter close/finish lifecycle. #69961
- Corrupted cache for PK SST tables is not clraered. #69693
AsyncFlushOutputStreamuse-after-free issue. #69688- Retention clock reset issue and incomplete scan in
disableRecoverPartitionWithSameName. #69677 - NPE in
StreamLoadMultiStmtTask.cancelAfterRestartafter deserialization. #69662 - Unnecessary RPCs and metadata queries caused by the incorrect logic of
SchemaBeTabletsScanner. #69645 - Graceful exit caused different transactions to publish the same version. #69639
TabletUpdates::get_column_valuescrashes with SIGSEGV when the Primary Key Index contains stale entries that point to rowsets that have been compacted. #69617KILL ANALYZEfails to stopANALYZE TABLEtasks. #69592- Unexpected behavior because not all exceptions of RowGroupWriter are caught. #69568
- TaskRun warehouse display issues after changing the warehouse for the materialized view. #69567
- Sort key does not include newly added key columns after schema change on aggregate and unique tables. #69529
- Issues caused by
isInternalCancelErrorusingequals. #69523 - TaskManager scheduling bugs after
ALTER MATERIALIZED VIEW. #69504 - Pipeline will be blocked or crash because not all exceptions of
ParquetFileWriter::closeare caught. #69492 - Materialized view force refresh bugs for partitioned tables. #69488
- Incorrect status was returned when certain writers failed to flush data. #69473
- Rowset files were deleted when Primary Key tablets were moved to trash. #69438
- INSERT failure when the range of an automatic partition is enclosed by an existing merged partition. #69429
- Materialized view tablet meta inconsistency between FE leader and follower. #69428
- Concurrency bugs related to function fields. #69315
- Premature deletion of data because rollup handler's active transaction ID is not considered in
computeMinActiveTxnId. #69285 - Lock leak in
addPartitionscaused by name-based table lookup after concurrent SWAP. #69284 DROP FUNCTION IF EXISTSignored theifExistsflag. #69216- Inconsistent behavior between StarRocks and MySQL-compatible syntax when
CAST(... AS SIGNED)in TypeParser. #69181 - Issue with case-insensitive partition lookup in query table copy. #69173
- Missing aggregate function when MIN/MAX stats rewrite failed. #69149
- CVE-2025-67721. #69138
- All-null value handling bug in synchronous materialized views. #69136
- Projection loss in materialized view rewrite due to shared mutable state. #69063
- Incorrect estimation of the Iceberg cache Weigher. #69058
FULL OUTER JOIN USINGissue with constant subqueries. #69028- Issue with
DISTINCT ORDER BYalias resolution for duplicated constants. #69014 - Issue when materialized view visits external catalog on reload. #68926
- NPE in Iceberg
getPartitions. #68907 - The container were not properly included in the Azure ABFS/WASB FileSystem cache key. #68901
- Erroneous query results after modifying
CHARcolumn length in shared-data clusters. #68808 - Case-insensitive issue with username in LDAP authentication. #67966
- Partitions could not be created after adding
storage_cooldown_ttlto a table. #60290
4.0.6β
Release Date: February 14, 2026
Improvementsβ
- Support Partition Transforms with parentheses when creating Iceberg tables (for example,
PARTITION BY (bucket(k1, 3))). #68945 - Removed the restriction that partition columns in Iceberg tables must be at the end of the column list; they can now be defined at any position. #68340
- Introduced host-level sorting for Iceberg table sink, controlled by the system variable
connector_sink_sort_scope(Default: FILE), to organize data layout for better read performance. #68121 - Improved error messages for Iceberg partition transform functions (for example,
bucket,truncate) when the argument count is incorrect. #68349 - Refactored table property handling to improve support for different file formats (ORC/Parquet) and compression codecs in Iceberg tables. #68588
- Added table-level query timeout configuration
table_query_timeoutfor fine-grained control (Priority: Session > Table > Cluster). #67547 - Supports the
ADMIN SHOW AUTOMATED CLUSTER SNAPSHOTstatement to view automated snapshot status and schedule. #68455 - Supports displaying the original user-defined SQL with comments in
SHOW CREATE VIEW. #68040 - Exposed Merge Commit-enabled Stream Load tasks in
information_schema.loadsfor better observability. #67879 - Introduced FE memory estimation utility API
/api/memory_usage. #68287 - Reduced unnecessary logging in
CatalogRecycleBinduring partition recycling. #68533 - Triggered refresh of related asynchronous materialized views when the base table undergoes Swap/Drop/Replace Partition operations. #68430
- Supports
VARBINARYtype forcount distinct-like aggregate functions. #68442 - Enhanced expression statistics to propagate histogram MCV for semantics-safe expressions (for example,
cast(k as bigint) + 10) to improve skew detection. #68292
Bug Fixesβ
The following issues have been fixed:
- Potential crashes in Skew Join V2 runtime filters. #67611
- Join predicate type mismatch (for example, INT = VARCHAR) caused by low-cardinality rewriting. #68568
- Issues in query queue allocation time and pending timeout logic. #65802
unique_idconflict for Flat JSON extended columns after schema changes. #68279- Concurrent partition access issues in
OlapTableSink.complete(). #68853 - Incorrect metadata tracking when restoring manually downloaded cluster snapshots. #68368
- Double slashes in backup paths when the repository location ends with
/. #68764 - OBS AK/SK credentials in the
SHOW CREATE CATALOGoutput were not masked. #65462
4.0.5β
Release Date: February 3, 2026
Improvementsβ
- Bumped Paimon version to 1.3.1. #67098
- Restored missing optimizations in DP statistics estimation to reduce redundant calculations. #67852
- Improved pruning in DP Join reorder to skip expensive candidate plans earlier. #67828
- Optimized JoinReorderDP partition enumeration to reduce object allocation and added an atom count cap (β€ 62). #67643
- Optimized DP join reorder pruning and added checks to BitSet to reduce stream operation overhead. #67644
- Skipped predicate column statistics collection during DP statistics estimation to reduce CPU overhead. #67663
- Optimized correlated Join row count estimation to avoid repeatedly building
Statisticsobjects. #67773 - Reduced memory allocations in
Statistics.getUsedColumns. #67786 - Avoided redundant
Statisticsmap copies when only row counts are updated. #67777 - Skipped aggregate pushdown logic when no aggregation exists in the query to reduce overhead. #67603
- Improved COUNT DISTINCT over windows, added support for fused multi-distinct aggregations, and optimized CTE generation. #67453
- Supports
map_aggfunction in the Trino dialect. #66673 - Supports batching retrieval of LakeTablet location information during physical planning to reduce RPC calls in shared-data clusters. #67325
- Added a thread pool for Publish Version transactions to shared-nothing clusters to improve concurrency. #67797
- Optimized LocalMetastore locking granularity by replacing database-level locks with table-level locks. #67658
- Refactored MergeCommitTask lifecycle management and added support for task cancellation. #67425
- Supports intervals for automated cluster snapshots. #67525
- Automatically cleaned up unused
mem_poolentries in MemTrackerManager. #67347 - Ignored
information_schemaqueries during warehouse idle checks. #67958 - Supports dynamically enabling global shuffle for Iceberg table sinks based on data distribution. #67442
- Added Profile metrics for connector sink modules. #67761
- Improved the collection and display of load spill metrics in Profiles, distinguishing between local and remote I/O. #67527
- Changed Async-Profiler log level to Error to avoid repeating warning logs. #67297
- Notified Starlet during BE shutdown to report SHUTDOWN status to StarMgr. #67461
Bug Fixesβ
The following issues have been fixed:
- Lacking support for legal simple paths containing hyphens (
-). #67988 - Runtime error when aggregate pushdown occurred on grouping keys involving JSON types. #68142
- Issue where JSON path rewrite rules incorrectly pruned partition columns referenced in partition predicates. #67986
- Type mismatch issue when rewriting simple aggregation using statistics. #67829
- Potential heap-buffer-overflow in partition Joins. #67435
- Duplicate
slot_idsintroduced when pushing down heavy expressions. #67477 - Division-by-zero error in ExecutionDAG fragment connection for lacking precondition checks. #67918
- Potential issues caused by fragment parallel prepare for single BE. #67798
- Operator terminates incorrectly for lacking
set_finishedmethod for RawValuesSourceOperator. #67609 - BE crash caused by unsupported DECIMAL256 type (precision > 38) in column aggregators. #68134
- Shared-data clusters lack support for Fast Schema Evolution v2 over DELETE operations by carrying
schema_keyin requests. #67456 - Shared-data clusters lack support for Fast Schema Evolution v2 over synchronous materialized views and traditional schema changes. #67443
- Vacuum might accidentally delete files when file bundling is disabled during FE downgrade. #67849
- Incorrect graceful exit handling in MySQLReadListener. #67917
4.0.4β
Release Date: January 16, 2026
Improvementsβ
- Supports Parallel Prepare for Operators and Drivers, and single-node batch fragment deployment to improve query scheduling performance. #63956
- Optimized
deltaRowscalculation with lazy evaluation for large partition tables. #66381 - Optimized Flat JSON processing with sequential iteration and improved path derivation. #66941 #66850
- Supports releasing Spill Operator memory earlier to reduce memory usage in group execution. #66669
- Optimized the logic to reduce string comparison overhead. #66570
- Improved skew detection in
GroupByCountDistinctDataSkewEliminateRuleandSkewJoinOptimizeRuleto support histogram and NULL-based strategies. #66640 #67100 - Enhanced Column ownership management in Chunk using Move semantics to reduce Copy-On-Write overhead. #66805
- For shared-data clusters, added FE
TableSchemaServiceand updatedMetaScanNodeto support Fast Schema Evolution v2 schema retrieval. #66142 #66970 - Supports multi-warehouse Backend resource statistics and parallelism (DOP) calculation for better resource isolation. #66632
- Supports configuring Iceberg split size via StarRocks session variable
connector_huge_file_size. #67044 - Supports label-formatted statistics in
QueryDumpDeserializer. #66656 - Added an FE configuration
lake_enable_fullvacuum(Default:false) to allow disabling Full Vacuum in shared-data clusters. #63859 - Upgraded lz4 dependency to v1.10.0. #67045
- Added fallback logic for sample-type cardinality estimation when row count is 0. #65599
- Validated Strict Weak Ordering property for lambda comparator in
array_sort. #66951 - Optimized error messages when fetching external table metadata (Delta/Hive/Hudi/Iceberg) fails, showing root causes. #66916
- Supports dumping pipeline status on query timeout and cancelling with
TIMEOUTstate in FE. #66540 - Displays matched rule index in SQL blacklist error messages. #66618
- Added labels to column statistics in
EXPLAINoutput. #65899 - Filtered out "cancel fragment" logs for normal query completions (for example, LIMIT reached). #66506
- Reduced Backend heartbeat failure logs when the warehouse is suspended. #66733
- Supports
IF EXISTSin theALTER STORAGE VOLUMEsyntax. #66691
Bug Fixβ
The following issues have been fixed:
- Incorrect
DISTINCTandGROUP BYresults under Low Cardinality optimization due to missingwithLocalShuffle. #66768 - Rewrite error for JSON v2 functions with Lambda expressions. #66550
- Incorrect application of Partition Join in Null-aware Left Anti Join within correlated subqueries. #67038
- Incorrect row count calculation in the Meta Scan rewrite rule. #66852
- Nullable property mismatched in Union Node when rewriting Meta Scan by statistics. #67051
- BE crash caused by optimization logic for Ranking window functions when
PARTITION BYandORDER BYare missing. #67094 - Potential wrong results in Group Execution Join with window functions. #66441
- Incorrect results from
PartitionColumnMinMaxRewriteRuleunder specific filter conditions. #66356 - Incorrect Nullable property deduction in Union operations after aggregation. #65429
- Crash in
percentile_approx_weightedwhen handling compression parameters. #64838 - Crash when spilling with large string encoding. #61495
- Crash triggered by multiple calls to
set_collectorwhen pushing down local TopN. #66199 - Dependency deduction error in LowCardinality rewrite logic. #66795
- Rowset ID leak when rowset commit fails. #66301
- Metacache lock contention. #66637
- Ingestion failure when column-mode partial update is used with conditional update. #66139
- Concurrent import failure caused by Tablet deletion during the ALTER operation. #65396
- Tablet metadata load error due to RocksDB iteration timeout. #65146
- Compression settings were not applied during table creation and Schema Change in shared-data clusters. #65673
- Delete Vector CRC32 compatibility issue during upgrade. #65442
- Status check logic error in file cleanup after clone task failure. #65709
- Abnormal statistics collection logic after
INSERT OVERWRITE. #65327 #65298 #65225 - Foreign Key constraints were lost after FE restart. #66474
- Metadata retrieval error after Warehouse deletion. #66436
- Inaccurate Audit Log scan statistics under high selectivity filters. #66280
- Incorrect query error rate metrics calculation logic. #65891
- Potential MySQL connection leaks when tasks exit. #66829
- BE status was not updated immediately on the SIGSEGV crash. #66212
- NPE during LDAP user login. #65843
- Inaccurate error log when switching users in HTTP SQL requests. #65371
- HTTP context leaks during TCP connection reuse. #65203
- Missing QueryDetail in Profile logs for queries forwarded from Follower. #64395
- Missing Prepare/Execute details in Audit logs. #65448
- Crash caused by HyperLogLog memory allocation failures. #66747
- Issue with the
trimfunction memory reservation. #66477 #66428 - CVE-2025-66566 and CVE-2025-12183. #66453 #66362 #67053
- Race condition in Exec Group driver submission. #66099
- Use-after-free risk in Pipeline countdown. #65940
MemoryScratchSinkOperatorhangs when the queue closes. #66041- Filesystem cache key collision issue. #65823
- Wrong subtask count in
SHOW PROC '/compactions'. #67209 - A unified JSON format is not returned in the Query Profile API. #67077
- Improper
getTableexception handling that affects the materialized view check. #67224 - Inconsistent output of the
Extracolumn from theDESCstatement for native and cloud-native tables. #67238 - Race condition in single-node deployments. #67215
- Log leakage from third-party libraries. #67129
- Incorrect REST Catalog authentication logic that causes authentication failures. #66861
4.0.3β
Release Date: December 25, 2025
Improvementsβ
- Supports
ORDER BYclauses for STRUCT data types #66035 - Supports creating Iceberg views with properties and displaying properties in the output of
SHOW CREATE VIEW. #65938 - Supports altering Iceberg table partition specs using
ALTER TABLE ADD/DROP PARTITION COLUMN. #65922 - Supports
COUNT/SUM/AVG(DISTINCT)aggregation over framed windows (for example,ORDER BY/PARTITION BY) with optimization options. #65815 - Optimized CSV parsing performance by using
memchrfor single-character delimiters. #63715 - Added an optimizer rule to push down Partial TopN to the Pre-Aggregation phase to reduce network overhead. #61497
- Enhanced Data Cache monitoring
- Optimized Sort and Aggregation operators to support rapid memory release in OOM scenarios. #66157
- Added
TableSchemaServicein FE for shared-data clusters to allow CNs to fetch specific schemas on demand. #66142 - Optimized Fast Schema Evolution to retain history schemas until all dependent ingestion jobs are finished. #65799
- Enhanced
filterPartitionsByTTLto properly handle NULL partition values to prevent all partitions from being filtered. #65923 - Optimized
FusedMultiDistinctStateto clear the associated MemPool upon reset. #66073 - Made
ICEBERG_CATALOG_SECURITYproperty check case-insensitive in Iceberg REST Catalog. #66028 - Added HTTP endpoint
GET /service_idto retrieve StarOS Service ID in shared-data clusters. #65816 - Replaced deprecated
metadata.broker.listwithbootstrap.serversin Kafka consumer configurations. #65437 - Added FE configuration
lake_enable_fullvacuum(Default: false) to allow disabling the Full Vacuum Daemon. #66685 - Updated lz4 library to v1.10.0. #67080
Bug Fixesβ
The following issues have been fixed:
latest_cached_tablet_metadatacould cause versions to be incorrectly skipped during batch Publish. #66558- Potential issues caused by
ClusterSnapshotrelative checks inCatalogRecycleBinwhen running in shared-nothing clusters. #66501 - BE crash when writing complex data types (ARRAY/MAP/STRUCT) to Iceberg tables during Spill operations. #66209
- Potential hang in Connector Chunk Sink when the writer's initialization or initial write fails. #65951
- Connector Chunk Sink bug where
PartitionChunkWriterinitialization failure caused a null pointer dereference during close. #66097 - Setting a non-existent system variable would silently succeed instead of reporting an error. #66022
- Bundle metadata parsing failure when Data Cache is corrupted. #66021
- MetaScan returned NULL instead of 0 for count columns when the result is empty. #66010
SHOW VERBOSE RESOURCE GROUP ALLdisplays NULL instead ofdefault_mem_poolfor resource groups created in earlier versions. #65982- A
RuntimeExceptionduring query execution after disabling theflat_jsontable configuration. #65921 - Type mismatch issue in shared-data clusters caused by rewriting
min/maxstatistics to MetaScan after Schema Change. #65911 - BE crash caused by ranking window optimization when
PARTITION BYandORDER BYare missing. #67093 - Incorrect
can_use_bfcheck when merging runtime filters, which could lead to wrong results or crashes. #67062 - Pushing down runtime bitset filters into nested OR predicates causes incorrect results. #67061
- Potential data race and data loss issues caused by write or flush operations after the DeltaWriter has finished. #66966
- Execution error caused by mismatched nullable properties when rewriting simple aggregation to MetaScan. #67068
- Incorrect row count calculation in the MetaScan rewrite rule. #66967
- Versions might be incorrectly skipped during batch Publish due to inconsistent cached tablet metadata. #66575
- Improper error handling for memory allocation failures in HyperLogLog operations. #66827
4.0.2β
Release Date: December 4, 2025
New Featuresβ
- Introduced a new resource group attribute,
mem_pool, allowing multiple resource groups to share the same memory pool and enforce a joint memory limit for the pool. This feature is backward compatible.default_mem_poolis used ifmem_poolis not specified. #64112
Improvementsβ
- Reduced remote storage access during Vacuum after File Bundling is enabled. #65793
- The File Bundling feature caches the latest tablet metadata. #65640
- Improved safety and stability for long-string scenarios. #65433 #65148
- Optimized the
SplitTopNAggregateRulelogic to avoid performance regression. #65478 - Applied the Iceberg/DeltaLake table statistics collection strategy to other external data sources to avoid collecting statistics when the table is a single table. #65430
- Added Page Cache metrics to the Data Cache HTTP API
api/datacache/app_stat. #65341 - Supports ORC file splitting to enable parallel scanning of a single large ORC file. #65188
- Added selectivity estimation for IF predicates in the optimizer. #64962
- Supports constant evaluation of
hour,minute, andsecondforDATEandDATETIMEtypes in the FE. #64953 - Enabled rewrite of simple aggregation to MetaScan by default. #64698
- Improved multiple-replica assignment handling in shared-data clusters for enhanced reliability. #64245
- Exposes cache hit ratio in audit logs and metrics. #63964
- Estimates per-bucket distinct counts for histograms using HyperLogLog or sampling to provide more accurate NDV for predicates and joins. #58516
- Supports FULL OUTER JOIN USING with SQL-standard semantics. #65122
- Prints memory information when Optimizer times out for diagnostics. #65206
Bug Fixesβ
The following issues have been fixed:
- DECIMAL56
mod-related issue. #65795 - Issue related to Iceberg scan range handling. #65658
- MetaScan rewrite issues on temporary partitions and random buckets. #65617
JsonPathRewriteRuleuses the wrong table after transparent materialized view rewrite. #65597- Materialized view refresh failures when
partition_retention_conditionreferenced generated columns. #65575 - Iceberg min/max value typing issue. #65551
- Issue with queries against
information_schema.tablesandviewsacross different databases whenenable_evaluate_schema_scan_ruleis set totrue. #65533 - Integer overflow in JSON array comparison. #64981
- MySQL Reader does not support SSL. #65291
- ARM build issue caused by SVE build incompatibility. #65268
- Queries based on bucket-aware execution may get stuck for bucketed Iceberg tables. #65261
- Robust error propagation and memory safety issues for the lack of memory limit checks in OLAP table scan. #65131
Behavior Changesβ
- When a materialized view is inactivated, the system recursively inactivates its dependent materialized views. #65317
- Uses the original materialized view query SQL (including comments/formatting) when generating SHOW CREATE output. #64318
4.0.1β
Release Date: November 17, 2025
Improvementsβ
- Optimized TaskRun session variable handling to process known variables only. #64150
- Supports collecting statistics of Iceberg and Delta Lake tables from metadata by default. #64140
- Supports collecting statistics of Iceberg tables with bucket and truncate partition transform. #64122
- Supports inspecting FE
/procprofile for debugging. #63954 - Enhanced OAuth2 and JWT authentication support for Iceberg REST catalogs. #63882
- Improved bundle tablet metadata validation and recovery handling. #63949
- Improved scan-range memory estimation logic. #64158
Bug Fixesβ
The following issues have been fixed:
- Transaction logs were deleted when publishing bundle tablets. #64030
- The join algorithm cannot guarantee the sort property because, after joining, the sort property is not reset. #64086
- Issues related to transparent materialized view rewrite. #63962
Behavior Changesβ
- Added the property
enable_iceberg_table_cacheto Iceberg Catalogs to optionally disable Iceberg table cache and allow it always to read the latest data. #64082 - Ensured
INSERT ... SELECTreads the freshest metadata by refreshing external tables before planning. #64026 - Increased lock table slots to 256 and added
ridto slow-lock logs. #63945 - Temporarily disabled
shared_scandue to incompatibility with event-based scheduling. #63543 - Changed the default Hive Catalog cache TTL to 24 hours and removed unused parameters. #63459
- Automatically determine the Partial Update mode based on the session variable and the number of inserted columns. #62091
4.0.0β
Release date: October 17, 2025
Data Lake Analyticsβ
- Unified Page Cache and Data Cache for BE metadata, and adopted an adaptive strategy for scaling. #61640
- Optimized metadata file parsing for Iceberg statistics to avoid repetitive parsing. #59955
- Optimized COUNT/MIN/MAX queries against Iceberg metadata by efficiently skipping over data file scans, significantly improving aggregation query performance on large partitioned tables and reducing resource consumption. #60385
- Supports compaction for Iceberg tables via procedure
rewrite_data_files. - Supports Iceberg tables with hidden partitions, including creating, writing, and reading the tables. #58914
- Supports setting sort keys when creating Iceberg tables.
- Optimizes sink performance for Iceberg tables.
- Iceberg Sink supports spilling large operators, global shuffle, and local sorting to optimize memory usage and address small file issues. #61963
- Iceberg Sink optimizes local sorting based on Spill Partition Writer to improve write efficiency. #62096
- Iceberg Sink supports global shuffle for partitions to further reduce small files. #62123
- Enhanced bucket-aware execution for Iceberg tables to improve concurrency and distribution capabilities of bucketed tables. #61756
- Supports the TIME data type in the Paimon catalog. #58292
- Upgraded Iceberg version to 1.10.0. #63667
Security and Authenticationβ
- In scenarios where JWT authentication and the Iceberg REST Catalog are used, StarRocks supports the passthrough of user login information to Iceberg via the REST Session Catalog for subsequent data access authentication. #59611 #58850
- Supports vended credentials for the Iceberg catalog.
- Supports granting StarRocks internal roles to external groups obtained via Group Provider. #63385 #63258
- Added REFRESH privilege to external tables to control the permission to refresh them. #63385
Storage Optimization and Cluster Managementβ
- Introduced β―the File Bundling optimization for the cloud-native table in shared-data clusters to automatically bundle the data files generated by loading, Compaction, or Publish operations, thereby reducing the API cost caused by high-frequency access to the external storage system. File Bundling is enabled by default for tables created in v4.0 or later. #58316
- Supports Multi-Table Write-Write Transaction to allow users to control the atomic submission of INSERT, UPDATE, and DELETE operations. The transaction supports Stream Load and INSERT INTO interfaces, effectively guaranteeing cross-table consistency in ETL and real-time write scenarios. #61362
- Supports Kafka 4.0 for Routine Load.
- Supports full-text inverted indexes on Primary Key tables in shared-nothing clusters.
- Supports modifying aggregate keys of Aggregate tables. #62253
- Supports enabling case-insensitive processing on names of catalogs, databases, tables, views, and materialized views. #61136
- Supports blacklisting Compute Nodes in shared-data clusters. #60830
- Supports global connection ID. #57256
- Added the
recyclebin_catalogsmetadata view to Information Schema to display recoverable deleted metadata. #51007
Query and Performance Improvementβ
- Supports DECIMAL256 data type, expanding the upper limit of precision from 38 to 76 bits. Its 256-bit storage provides better adaptability to high-precision financial and scientific computing scenarios, effectively mitigating DECIMAL128's precision overflow problem in very large aggregations and high-order operations. #59645
- Improved the performance for basic operators.#61691 #61632 #62585 #61405 #61429
- Optimized the performance of the JOIN and AGG operators. #61691
- [Preview] Introduced SQL Plan Manager to allow users to bind a query plan to a query, thereby preventing the query plan from changing due to system state changes (mainly data updates and statistics updates), thus stabilizing query performance. #56310
- Introduced Partition-wise Spillable Aggregate/Distinct operators to replace the original Spill implementation based on sorted aggregation, significantly improving aggregation performance and reducing read/write overhead in complex and high-cardinality GROUP BY scenarios. #60216
- Flat JSON V2:
- Supports configuring Flat JSON on the table level. #57379
- Enhance JSON columnar storage by retaining the V1 mechanism while adding page- and segment-level indexes (ZoneMaps, Bloom filters), predicate pushdown with late materialization, dictionary encoding, and integration of a low-cardinality global dictionary to significantly boost execution efficiency. #60953
- Supports an adaptive ZoneMap index creation strategy for the STRING data type. #61960
- Enhanced query observability:
- Optimized EXPLAIN ANALYZE output to display the execution metrics by group and by operator for better readability. #63326
QueryDetailActionV2andQueryProfileActionV2now support JSON format, enhancing cross-FE query capabilities. #63235- Supports retrieving Query Profile information across all FEs. #61345
- SHOW PROCESSLIST statements display Catalog, Query ID, and other information. #62552
- Enhanced query queue and process monitoring, supporting display of Running/Pending statuses.#62261
- Materialized view rewrites consider the distribution and sort keys of the original table, improving the selection of optimal materialized views. #62830
Functions and SQL Syntaxβ
- Added the following functions:
- Provides the following syntactic extensions:
Behavior Changesβ
- Adjust the logic of the materialized view parameter
auto_partition_refresh_numberto limit the number of partitions to refresh regardless of auto refresh or manual refresh. #62301 - Flat JSON is enabled by default. #62097
- The default value of the system variable
enable_materialized_view_agg_pushdown_rewriteis set totrue, indicating that aggregation pushdown for materialized view query rewrite is enabled by default. #60976 - Changed the type of some columns in
information_schema.materialized_viewsto better align with the corresponding data. #60054 - The
split_partfunction returns NULL when the delimiter is not matched. #56967 - Use STRING to replace fixed-length CHAR in CTAS/CREATE MATERIALIZED VIEW to avoid deducing the wrong column length, which may cause materialized view refresh failures. #63114 #62476
- Data Cache-related configurations are simplified. #61640
datacache_mem_sizeanddatacache_disk_sizeare now effective.storage_page_cache_limit,block_cache_mem_size,block_cache_disk_sizeare deprecated.
- Added new catalog properties (
remote_file_cache_memory_ratiofor Hive, andiceberg_data_file_cache_memory_usage_ratioandiceberg_delete_file_cache_memory_usage_ratiofor Iceberg) to limit the memory resources used for Hive and Iceberg metadata cache, and set the default values to0.1(10%). Adjust the metadata cache TTL to 24 hours. #63459 #63373 #61966 #62288 - SHOW DATA DISTRIBUTION now will not merge the statistics of all materialized indexes with the same bucket sequence number. It only shows data distribution at the materialized index level. #59656
- The default bucket size for automatic bucket tables is changed from 4GB to 1GB to improve performance and resource utilization. #63168
- The system determines the Partial Update mode based on the corresponding session variable and the number of columns in the INSERT statement. #62091
- Optimized the
fe_tablet_schedulesview in the Information Schema. #62073 #59813- Renamed the
TABLET_STATUScolumn toSCHEDULE_REASON, theCLONE_SRCcolumn toSRC_BE_ID, and theCLONE_DESTcolumn toDEST_BE_ID. - The data types of the
CREATE_TIME,SCHEDULE_TIMEandFINISH_TIMEcolumns have been changed fromDOUBLEtoDATETIME.
- Renamed the
- The
is_leaderlabel has been added to some FE metrics. #63004 - Shared-data clusters using Microsoft Azure Blob Storage and Data Lake Storage Gen 2 as object storage will experience Data Cache failure after being upgraded to v4.0. The system will automatically reload the cache.