SynapCores v2.1.0-ce — Release Notes
v2.1 delivers DuckDB-class analytics on your own S3 bucket, flat memory under sustained load, high-throughput bulk ingestion, and a real ONNX inference runtime — all in the same single binary.
New Features
Parallel analytics engine — DuckDB-class speed on your own bucket
Grouped aggregation now fans out across every core and merges partials;
dictionary-encoded Parquet columns are aggregated by code instead of being
expanded to strings; COUNT(DISTINCT) rides the same parallel path.
On a 13-query suite over identical S3 Parquet objects, v2.1.0 runs at a
median 1.08× of DuckDB 1.5.5 — parity — and wins outright on 4 of 13
shapes, including a two-column GROUP BY over 230k rows at 1.9× faster
(2.7 ms vs 5.2 ms). An external join + GROUP BY across 189k × 28k rows
answers in 13 ms — dead even with DuckDB.
Business benefit: interactive dashboards and exploratory joins directly on
the Parquet you already have in S3 — no warehouse, no copy step, no second
vendor. Tune with AIDB_AGGREGATE_WORKERS (optional; defaults are right for
most hosts).
Warm S3 metadata and block cache — one round trip per query
Foreign-table queries previously re-paid ~80 object-store round trips each.
v2.1 keeps an ETag-validated metadata cache and a block cache
(AIDB_LAKE_DATA_CACHE_MB, default 64), bringing a warm query down to one
round trip. Changed objects are detected by ETag, so the cache never serves
a replaced file's bytes.
Business benefit: on a real S3 endpoint this removes 0.8–1.6 seconds of pure network latency from every lake query — the difference between a dashboard that feels live and one that feels remote.
Flat memory under sustained load
The engine now runs on jemalloc, and every memory-heavy operator — joins,
sorts, aggregates, result materialization — charges a per-query budget
(256 MiB default, AIDB_QUERY_MEMORY_BYTES). A query that exceeds its budget
returns a clear error; the engine keeps serving everyone else.
Proof run on the release build: 80 consecutive heavy aggregation queries, including 40 over a 5M-row table — RSS moved ±12 MiB.
Business benefit: you can put SynapCores behind a refreshing dashboard and leave it running. Memory is predictable, capacity planning is possible, and one oversized query can no longer take the process down.
High-throughput bulk load
The native write path batches documents into bounded chunks, streams uploads, and retains verified indexes through the load.
On the release build: a 1,000,000-row CSV imports at ~20,000 rows/s
sustained, with exact row counts and checksummed aggregates on arrival;
batched INSERT sustains ~25,000 rows/s.
Business benefit: million-row migrations and nightly loads complete in minutes over the standard API — no special tooling, no degradation as the table grows.
MySQL-style type coercion in comparisons and joins
Comparing a numeric column against a text literal now matches numerically:
WHERE int_col = '06' finds 6, and an INT column joins a VARCHAR column by
value ('0123' joins 123). Same-type comparisons are unchanged, and index
probes stay plan-consistent.
Business benefit: data bulk-loaded from CSV or migrated from MySQL "just works" with the queries written against it — no silent empty results from invisible type mismatches.
ONNX Runtime inference backend
The engine now embeds the ONNX Runtime (CPU execution provider) with pinned
NLI cross-encoder models and semantic verification recipes. A new
--onnx-self-test flag runs input-dependent inference cases through the
packaged runtime and reports the backend:
{"backend":"onnxruntime","provider":"CPU","input_dependent":true,"cases":2}
Business benefit: semantic verification and cross-encoder scoring run in-process, next to the data — no Python sidecar, no model-serving infrastructure to operate.
Export the engine's own tables to your data lake
Lake export jobs now accept source_kind=local: any engine table can be
exported to Parquet on S3 on a schedule, with read-back verification —
previously only external MySQL sources could be exported. Jobs now also
report a live next_run_at.
Business benefit: one instance can archive its own hot tables to cheap object storage and keep querying them in place — tiered storage without an ETL tool.
AutoML observability
automl_experiments, automl_models and automl_trials are queryable SQL
catalogs; SHOW MODELS lists deployed models; max_trials drives a real
hyperparameter search; training reports score_basis, failed_trials,
trials_attempted and ignored_options; DECIMAL columns are accepted as
features.
Business benefit: model training is auditable with plain SQL — what was tried, what failed, and what's deployed, without leaving the database.
Bug Fixes
SQL engine
GROUP BYaccepts column aliases and ordinals (GROUP BY 1,GROUP BY region_name).- CTE declared column lists are applied:
WITH c(n) AS (SELECT 1) SELECT n FROM cworks. - Recursive CTE termination predicates are preserved; constant-false recursion terminates immediately.
- A
GROUP BYwhose keys overflowed the hash table returned zero rows with no error; it now returns the error. EXTRACT(...)and the date/time function family evaluate inside joinONclauses.max_rowsbounds query execution, not just response serialization.- Ordered-index selection under
LIMITno longer overrides a better predicate index.
Joins and lake
- An equi-join between two external Parquet tables ran as a nested loop (5.2 billion comparisons); it now uses the hash join and answers in 24 ms.
- Projecting columns from an external-table join allocated memory without bound; fixed by per-side projection pushdown.
- A mixed INT/VARCHAR equi-join returned 0 rows silently; keys now join numerically.
COUNT(*)after a ~1M-row import failed with a budget error; scan leases are now released per batch.- A
GROUP BYon a lake table retained ~150 MB per execution; memory is now flat.
DDL and catalogs
CREATE TABLEover an existing name returned "created successfully" without creating anything; it now returns an error.CREATE TABLE IF NOT EXISTSis unchanged.CREATE MEMORYbacking tables were recorded under the wrong tenant scope and unreachable from SQL.- The schema API reported
is_primary=falsefor every column; primary keys and their indexes are now reported. - AutoML training failed with "Table or collection does not exist" under tenant-scoped storage.
- The AutoML experiments catalog exposes both
idandexperiment_id.
Installation and operations
- Native installers now generate
AIDB_ENCRYPTION_KEY;/v1/lake/*no longer answers 503 on a fresh install. - Native installer HTTP timeout raised to 300 s to match Docker (cold
GENERATEwas cut off at 30 s). - Query responses serialize directly from engine values (~1.7 ms saved on a 3,226-row response, byte-identical wire format).
Upgrading from v2.0.0
- Duplicate
CREATE TABLEnow errors. Scripts that relied on the prior silent no-op should useCREATE TABLE IF NOT EXISTS. - Queries can fail closed on memory. Raise
AIDB_QUERY_MEMORY_BYTESfor known-heavy workloads. - Numeric/text comparisons coerce. If you depended on type mismatches returning no rows, compare as TEXT explicitly.
- New optional tuning knobs:
AIDB_AGGREGATE_WORKERS,AIDB_LAKE_DATA_CACHE_MB,AIDB_QUERY_MEMORY_BYTES.
Validation shipped with this tag: 245/245 state-asserting feature validations, full recipe-certification sweep with zero regressions against v2.0.0, the 13-query DuckDB benchmark protocol, a 1M-row bulk-load gate, and a non-AVX-512 hardware canary running the published artifact end to end, including the ONNX self-test.
[Download v2.1.0-ce →] · Docker: synapcores/community:v2.1.0-ce