Natural Language to SQL: How It Works and Why It Gets Numbers Wrong

Natural language to SQL, also called text-to-SQL or NL2SQL, lets you ask a database a question in plain English and get an answer back. When it works, it removes the wait for an analyst. When it fails, it fails quietly: the query runs, a number appears, and the number is wrong. This guide explains how the technology works, the six ways it goes wrong, and what to check before you trust it.

Trinetro Labs · Published 2 October 2026 · 5 min read

What natural language to SQL actually does

A database only understands SQL. A natural language to SQL system sits between you and the database and does three things.

  1. Reads the schema. It looks at the tables and columns in your database to learn what exists.
  2. Writes a query. It turns your question into a SQL statement. Most tools use a large language model for this step; others use rules.
  3. Runs it and shows the result. The SQL runs on your database and the rows come back as a table, a chart or a sentence.

The risk lives in the second step. A model that predicts likely SQL is not the same as a system that knows what your data means.

Six ways it returns wrong numbers

1. Double counting across tables

Orders and order lines live in separate tables. If a query joins them and then adds up the order total, each order is counted once per line. A ₹1,000 order with three lines becomes ₹3,000.

-- Valid SQL, wrong answer: each order is counted once per line item
SELECT SUM(o.order_total)
FROM orders o
JOIN order_items i ON i.order_id = o.id;

The query runs without an error. Only something that knows the grain of each table, one row per order versus one row per line, can see the mistake.

2. Guessing which column you meant

"Revenue" might be gross sales, net of discounts, net of returns, or booked versus invoiced. A tool that picks one silently gives you a number you cannot reconcile.

3. Vague time periods

"Last quarter" differs between a calendar year and an April to March financial year, which most Indian companies use. "Last month" on the first of the month can mean a day or a whole month.

4. Invented columns and tables

Language models sometimes produce a column that does not exist. Usually that causes an error. Occasionally it maps to a similarly named column and quietly returns the wrong data.

5. Different answers on different days

The same question can produce different SQL on different runs. Ask on Monday and on Friday and the totals may not match, even though the data did not change.

6. No way to check

If the tool shows only a number, you cannot tell which of these went wrong. Trust requires being able to inspect the tables used, the joins made and the metric definition.

Probabilistic and deterministic approaches

Language-model SQL generation compared with a deterministic engine
What mattersLanguage-model SQL generationDeterministic engine
How the query is builtPredicted by a model from the schema and the questionConstructed by rules from the detected schema, grain and joins
Same question twiceMay differIdentical
Joins across tablesChosen by the model; double counting is possibleScored and validated; grain is detected first
Metric definitionsGuessed from column namesResolved from a governed definition
Shows its workingSometimesAlways, as a reasoning trace
When it is unsureA plausible wrong numberA clarifying question or a stated limit

A six-point checklist before you trust one

  1. Ask the same question three times and compare the answers.
  2. Ask for a total you can verify somewhere else.
  3. Ask a question that crosses two tables and check it against a manual count.
  4. Use your own wording for a metric, such as your definition of an active customer, and see whether it asks or guesses.
  5. Ask for last quarter and check which dates it used.
  6. Find where the tool shows the query, the tables and the definitions behind each answer.

How Trinetro approaches it

Trinetro Labs treats natural language to SQL as a reasoning problem, not a prediction problem. It profiles your schema, detects the grain of each table, scores joins and resolves each business term to a definition before any query is built. The answer path contains no language model, so the same question on the same data returns the same result, and every answer carries a reasoning trace you can read.

Its governed team memory keeps the definitions. Once an admin approves what "revenue" means for your company, the question is not asked again and everyone gets the same number.

Frequently asked questions

What is natural language to SQL?

It is technology that converts a plain-English question into a SQL query, runs it on a database and returns the result, so people who do not write SQL can query data directly. It is also called text-to-SQL or NL2SQL.

Is text-to-SQL accurate?

It can be, but accuracy depends on the design. Language-model-based tools can return plausible wrong numbers, for example by double counting across joined tables. Test any tool on questions where you already know the answer, and prefer tools that show how each answer was produced.

Can ChatGPT write SQL for my database?

It can draft SQL if you paste in your schema, but it does not know your data, your definitions or your financial year, and the SQL it writes can differ between runs. Treat the output as a draft to review, not a verified answer.

What is the difference between NL2SQL and a semantic layer?

A semantic layer defines your metrics and relationships once, so every query uses the same meaning. Good NL2SQL tools either have one or build one automatically. Without it, the tool has to guess.

Do I need to know SQL to use natural language to SQL tools?

No. That is the point of the technology. It still helps to know which question you are asking and which numbers you can verify elsewhere.

Keep reading

Start free for 30 days · Book a demo · Try the guided demo