SynapCores v2.1 — An engine you can leave running, at DuckDB pace

Published on October 11, 2026

SynapCores v2.1.0-ce — Release Notes

v2.1 delivers DuckDB-class analytics on your own S3 bucket, flat memory under sustained load, high-throughput bulk ingestion, and a real ONNX inference runtime — all in the same single binary.


New Features

Parallel analytics engine — DuckDB-class speed on your own bucket

Grouped aggregation now fans out across every core and merges partials; dictionary-encoded Parquet columns are aggregated by code instead of being expanded to strings; COUNT(DISTINCT) rides the same parallel path.

On a 13-query suite over identical S3 Parquet objects, v2.1.0 runs at a median 1.08× of DuckDB 1.5.5 — parity — and wins outright on 4 of 13 shapes, including a two-column GROUP BY over 230k rows at 1.9× faster (2.7 ms vs 5.2 ms). An external join + GROUP BY across 189k × 28k rows answers in 13 ms — dead even with DuckDB.

Business benefit: interactive dashboards and exploratory joins directly on the Parquet you already have in S3 — no warehouse, no copy step, no second vendor. Tune with AIDB_AGGREGATE_WORKERS (optional; defaults are right for most hosts).

Warm S3 metadata and block cache — one round trip per query

Foreign-table queries previously re-paid ~80 object-store round trips each. v2.1 keeps an ETag-validated metadata cache and a block cache (AIDB_LAKE_DATA_CACHE_MB, default 64), bringing a warm query down to one round trip. Changed objects are detected by ETag, so the cache never serves a replaced file's bytes.

Business benefit: on a real S3 endpoint this removes 0.8–1.6 seconds of pure network latency from every lake query — the difference between a dashboard that feels live and one that feels remote.

Flat memory under sustained load

The engine now runs on jemalloc, and every memory-heavy operator — joins, sorts, aggregates, result materialization — charges a per-query budget (256 MiB default, AIDB_QUERY_MEMORY_BYTES). A query that exceeds its budget returns a clear error; the engine keeps serving everyone else.

Proof run on the release build: 80 consecutive heavy aggregation queries, including 40 over a 5M-row table — RSS moved ±12 MiB.

Business benefit: you can put SynapCores behind a refreshing dashboard and leave it running. Memory is predictable, capacity planning is possible, and one oversized query can no longer take the process down.

High-throughput bulk load

The native write path batches documents into bounded chunks, streams uploads, and retains verified indexes through the load.

On the release build: a 1,000,000-row CSV imports at ~20,000 rows/s sustained, with exact row counts and checksummed aggregates on arrival; batched INSERT sustains ~25,000 rows/s.

Business benefit: million-row migrations and nightly loads complete in minutes over the standard API — no special tooling, no degradation as the table grows.

MySQL-style type coercion in comparisons and joins

Comparing a numeric column against a text literal now matches numerically: WHERE int_col = '06' finds 6, and an INT column joins a VARCHAR column by value ('0123' joins 123). Same-type comparisons are unchanged, and index probes stay plan-consistent.

Business benefit: data bulk-loaded from CSV or migrated from MySQL "just works" with the queries written against it — no silent empty results from invisible type mismatches.

ONNX Runtime inference backend

The engine now embeds the ONNX Runtime (CPU execution provider) with pinned NLI cross-encoder models and semantic verification recipes. A new --onnx-self-test flag runs input-dependent inference cases through the packaged runtime and reports the backend:

{"backend":"onnxruntime","provider":"CPU","input_dependent":true,"cases":2}

Business benefit: semantic verification and cross-encoder scoring run in-process, next to the data — no Python sidecar, no model-serving infrastructure to operate.

Export the engine's own tables to your data lake

Lake export jobs now accept source_kind=local: any engine table can be exported to Parquet on S3 on a schedule, with read-back verification — previously only external MySQL sources could be exported. Jobs now also report a live next_run_at.

Business benefit: one instance can archive its own hot tables to cheap object storage and keep querying them in place — tiered storage without an ETL tool.

AutoML observability

automl_experiments, automl_models and automl_trials are queryable SQL catalogs; SHOW MODELS lists deployed models; max_trials drives a real hyperparameter search; training reports score_basis, failed_trials, trials_attempted and ignored_options; DECIMAL columns are accepted as features.

Business benefit: model training is auditable with plain SQL — what was tried, what failed, and what's deployed, without leaving the database.


Bug Fixes

SQL engine

  • GROUP BY accepts column aliases and ordinals (GROUP BY 1, GROUP BY region_name).
  • CTE declared column lists are applied: WITH c(n) AS (SELECT 1) SELECT n FROM c works.
  • Recursive CTE termination predicates are preserved; constant-false recursion terminates immediately.
  • A GROUP BY whose keys overflowed the hash table returned zero rows with no error; it now returns the error.
  • EXTRACT(...) and the date/time function family evaluate inside join ON clauses.
  • max_rows bounds query execution, not just response serialization.
  • Ordered-index selection under LIMIT no longer overrides a better predicate index.

Joins and lake

  • An equi-join between two external Parquet tables ran as a nested loop (5.2 billion comparisons); it now uses the hash join and answers in 24 ms.
  • Projecting columns from an external-table join allocated memory without bound; fixed by per-side projection pushdown.
  • A mixed INT/VARCHAR equi-join returned 0 rows silently; keys now join numerically.
  • COUNT(*) after a ~1M-row import failed with a budget error; scan leases are now released per batch.
  • A GROUP BY on a lake table retained ~150 MB per execution; memory is now flat.

DDL and catalogs

  • CREATE TABLE over an existing name returned "created successfully" without creating anything; it now returns an error. CREATE TABLE IF NOT EXISTS is unchanged.
  • CREATE MEMORY backing tables were recorded under the wrong tenant scope and unreachable from SQL.
  • The schema API reported is_primary=false for every column; primary keys and their indexes are now reported.
  • AutoML training failed with "Table or collection does not exist" under tenant-scoped storage.
  • The AutoML experiments catalog exposes both id and experiment_id.

Installation and operations

  • Native installers now generate AIDB_ENCRYPTION_KEY; /v1/lake/* no longer answers 503 on a fresh install.
  • Native installer HTTP timeout raised to 300 s to match Docker (cold GENERATE was cut off at 30 s).
  • Query responses serialize directly from engine values (~1.7 ms saved on a 3,226-row response, byte-identical wire format).

Upgrading from v2.0.0

  1. Duplicate CREATE TABLE now errors. Scripts that relied on the prior silent no-op should use CREATE TABLE IF NOT EXISTS.
  2. Queries can fail closed on memory. Raise AIDB_QUERY_MEMORY_BYTES for known-heavy workloads.
  3. Numeric/text comparisons coerce. If you depended on type mismatches returning no rows, compare as TEXT explicitly.
  4. New optional tuning knobs: AIDB_AGGREGATE_WORKERS, AIDB_LAKE_DATA_CACHE_MB, AIDB_QUERY_MEMORY_BYTES.

Validation shipped with this tag: 245/245 state-asserting feature validations, full recipe-certification sweep with zero regressions against v2.0.0, the 13-query DuckDB benchmark protocol, a 1M-row bulk-load gate, and a non-AVX-512 hardware canary running the published artifact end to end, including the ONNX self-test.

[Download v2.1.0-ce →] · Docker: synapcores/community:v2.1.0-ce