Replace Fivetran + Snowflake: MySQL to S3 Parquet Analytics, Self-Hosted

Published on October 10, 2026

Here is the analytics stack most teams end up with, usually without ever deciding on it:

  1. The application writes to MySQL or PostgreSQL.
  2. An ETL/ELT tool — Fivetran, Airbyte Cloud, Stitch — copies those tables into a warehouse, billed by rows moved.
  3. Snowflake (or BigQuery, or Redshift) stores a second copy and bills for every minute a warehouse is running queries.
  4. Dashboards and analysts query the warehouse.

It works. It is also two vendors, two bills that both grow with your data, and a copy of your production data living on someone else's platform — all to answer questions about data you already own.

This post shows a shorter path: export the operational tables to Parquet in your own S3 bucket on a schedule, and query that Parquet in place with plain SQL. One self-hosted engine does the export, the scheduling, the verification, and the querying. The engine is SynapCores; the Community Edition is free.

I'll be upfront about where this fits and where it doesn't — the honest version is at the end, and it matters.

What you are actually paying for

The ETL tool charges for moving rows. Fivetran, for example, bills on monthly active rows — every row inserted, updated, or deleted in a synced table counts. Busy operational tables (orders, events, sessions) are exactly the ones that churn the most rows, so the bill tracks your growth.

The warehouse charges for compute time. On Snowflake, a virtual warehouse burns credits for every second it is running, with a 60-second minimum each time it resumes. An X-Small warehouse uses 1 credit per hour, a Small uses 2, and each size up doubles. Dashboards that refresh every few minutes keep a warehouse awake most of the day.

Storage is the cheap part on both sides. The money is in moving the data and in running queries over it.

The alternative: MySQL → Parquet in your S3 → SQL, in one engine

SynapCores v2.0 added a data-lake export built into the database:

  • A connection to your MySQL holds the credentials once.
  • A destination points at your S3-compatible bucket (AWS S3, MinIO, Cloudflare R2, Wasabi). The secret is encrypted at rest and never returned by the API.
  • A job exports tables to Parquet on a cron schedule you set — a full snapshot per run for mutable tables like orders and customers, or incremental daily partitions for append-only tables like events and logs.
  • Each run reports rows, bytes, peak memory, and exit code, and verified_partitions confirms every file the job claims to have written was read back. No silent half-runs at 3am.
  • The Parquet is then registered as an external SQL table and queried in place. No load step, no second copy inside a warehouse.

The end state looks like this:

-- Parquet in s3://your-bucket/..., exported nightly from MySQL
SELECT region, COUNT(*) AS orders, SUM(amount) AS revenue
FROM orders_snap
WHERE status = 'shipped' AND snapshot_dt = '2026-10-09'
GROUP BY region
ORDER BY revenue DESC;

snapshot_dt is a real partition column, so filtering on it skips whole directories in S3 instead of reading them. Amounts come back as exact DECIMALs, and the recipe walks through checking the totals against MySQL to the cent.

The full, step-by-step setup — engine config, connection, destination, job, run, register, query — is in the recipe: Move a MySQL Table to S3 Parquet, Then Query It with SQL. It takes about 20 minutes.

Is it fast enough to replace a warehouse for this?

For the shape of work most teams send to a warehouse from their operational data — filters, scans, joins, rollups for dashboards — querying Parquet in place is fast. In the v2.0.0 release we benchmarked SynapCores against DuckDB on the same Parquet objects in the same bucket:

Query SynapCores v2.0 DuckDB 1.5.5
Scan 189,361 rows of a lake table 10.1 ms 28.6 ms
Count a join across two lake tables 5.5 ms 10.9 ms
Join returning 176,400 rows 16.5 ms 29.1 ms
GROUP BY one text column 5.0 ms 3.8 ms

We publish the row we lose: wide GROUP BY aggregation is still faster on DuckDB. For dashboards and exploratory joins, millisecond answers over your own bucket are realistic.

Because SynapCores speaks the MySQL wire protocol, many BI tools and drivers that already talk to MySQL can connect to it directly — test yours before you switch.

A worked cost example

Prices vary by contract, region, and edition, so treat this as a method, not a quote. Plug in your own invoice numbers.

Today (ETL + warehouse):

  • Warehouse compute: a Small warehouse (2 credits/hour) kept awake ~10 hours a day by dashboards and analysts = 20 credits/day ≈ 600 credits/month. At a typical on-demand rate of $2–$4 per credit, that is $1,200–$2,400/month — before any bigger warehouse for month-end reports.
  • ETL: whatever your row-based bill is. Look at your monthly active rows for the busiest tables; it is usually the line item that grows fastest.

With SynapCores CE on your own infrastructure:

  • Software: $0 — the Community Edition is free.
  • Server: one VM sized for your query load, e.g. 8 vCPU / 32 GB on-demand in a public cloud, roughly $250–$300/month, less with reserved pricing or hardware you already run.
  • S3 storage: about $23 per TB-month on S3 Standard. Parquet is compressed and columnar, so operational tables usually shrink considerably.
  • Data movement: the export reads from MySQL and writes to your bucket. Keep the engine in the same region as the bucket to avoid transfer charges.

The structural difference is that neither of the big variable costs — rows moved and warehouse minutes — exists anymore. Your cost becomes one server plus storage, and it stops scaling with how often people refresh a dashboard.

Where this fits — and where it doesn't

This is the part to read before you cancel anything.

It fits well when:

  • Your analytics sources are your own operational databases (the MySQL path is the one we have validated end to end).
  • Nightly or scheduled freshness is fine — reporting, finance, ops dashboards, analyst exploration.
  • A small-to-medium number of people query the data, not hundreds of concurrent users.
  • You want the data in an open format in your own bucket, readable by any Parquet tool, not locked in a vendor's storage.

Keep the warehouse (or the ETL tool) when:

  • You need near-real-time replication with change data capture. The export runs on a schedule: snapshot copies the full table per run, and incremental mode is for append-only tables — it does not replay updates and deletes.
  • You pull from dozens of SaaS sources (Salesforce, HubSpot, Stripe, ad platforms). Connector breadth is what ETL vendors are genuinely good at.
  • You need high concurrency, data sharing, or fine-grained governance across a large organization — that is what a cloud warehouse is built for.

Plenty of teams land in the middle: move the operational-database analytics — usually the biggest and most predictable part of the bill — to S3 + SynapCores, and keep the warehouse for the cross-source, high-concurrency work. That alone removes most of the rows you pay to move and most of the minutes you pay to query.

Try it

  1. Download the free Community Edition — macOS, Linux, or Docker.
  2. Follow the recipe: MySQL table to S3 Parquet, then query it with SQL.
  3. Point one dashboard at the external table and compare the numbers — and the bill — for a month.

Related reading: AI-Native Database vs Snowflake · SynapCores v2.0.0 release notes · Best AI-native databases in 2026