Five minutes, no pipeline, no warehouse. Connect a collection, let the schema be counted rather than guessed, then write ordinary SQL — including JOINs to a relational database.
Add a source, pick MongoDB, and give it a host, port and database. A user is optional — local instances usually run without auth.
Varan reads every document and counts the field set, the type histogram, presence and array element types — including nested paths. It is a census, not a sample, so nothing rare is invisible.
The counting happens on your machine: Varan walks the raw BSON locally rather than asking mongod to parse and aggregate. Measured with MongoDB's own profiler on 192,387 documents, that took the server's cost for a full census from 16,217 ms of CPU down to 195 ms — 83× less — for a byte-identical result across all 950 field paths. The collection is read 1.01× instead of 5.13×, and the read cursor is held open for a sixth as long.
The sidebar labels each collection exact or sampled next to its column count, so an incomplete schema can never be mistaken for a complete one. You can also ask per query: prefix it with EXACT to force a census, or FAST to skip one.
Each collection becomes a table. Nested fields become dotted columns:
orders
_id VARCHAR
status VARCHAR -- enum: new, paid, shipped
amount VARCHAR -- int x4850, string x150 -> needs a decision
addr.city VARCHAR
addr.geo.lat VARCHAR
tags VARCHAR[] -- native list: unnest(tags) works
items VARCHAR -- array of documents, held as JSON
Open schema mapping on the collection. It lists only what could not be settled by counting — typically a handful of fields — sorted so the least certain is first, each showing the evidence:
amount type → string destructive
across the whole collection: int ×4850, string ×150 — 3.0% minority, held as text until you decide
Pick the right answer and it is remembered. Re-scan later and your decision survives; only genuinely new fields come back as questions.
Decisions are tiered by what being wrong would cost. Cosmetic ones apply silently when the evidence is overwhelming. Anything that would change the shape of the table, or change what gets written back to your database, always asks.
SELECT c.name,
c."addr.city" AS city,
COUNT(o.id) AS orders,
SUM(o.total::DECIMAL) AS revenue
FROM mongo_customers c -- MongoDB
JOIN postgres_orders o ON o.customer_id = c._id -- PostgreSQL
GROUP BY 1, 2
ORDER BY revenue DESC;
Arrays behave like arrays:
SELECT tag, COUNT(*)
FROM (SELECT unnest(tags) AS tag FROM mongo_orders)
GROUP BY 1 ORDER BY 2 DESC;
Edits to a MongoDB table stay local until you say otherwise. Fixing a collection is iterative — you correct a column, look at it, correct it again — so pushing every intermediate statement would be a trap. When you are ready:
APPLY mongo_orders;
The write goes back as a per-document patch keyed on _id — only the fields you actually changed, as $set on individual paths. That matters: because the column list comes from what was discovered, replacing whole documents would erase anything not modelled. Patching cannot.
APPLY mongo_orders TYPES (amount AS STRING). Across 492,000 real documents, exactly one column needed a human answer.DRY RUN to see precisely what would be written without writing it.Relational sources behave differently on purpose: their schemas are declared, so their types are not in question. An UPDATE against a MySQL, MariaDB or PostgreSQL table writes straight through as it runs. SQL Server, Oracle and SQLite do not have a write-back path yet — there, Varan changes the local copy and says so explicitly, naming the table it did not write to.
It is worth being straight about the boundaries: