Luca Morelli
Software analyst and developer
Team lead

Applications with built-in AI: the model where it helps, code where exactness matters.

The most complex case I have worked on so far is a hybrid RAG system: an application that, starting from imported delivery documents, lets you query business data in natural language. It is the classic "chat with your data", with one requirement that changes everything: precision.

The figure must matchThe one in the management system, not just resemble it.
Work in progressA growing skill, not an established specialisation: the technologies are changing quickly.
RAGPythonPostgreSQLQdrantOpen models

The problem: precision

If you ask which products sold best over a period, the result must match what the management system says.

A classic vector store, with semantic search alone, cannot guarantee that precision. You need both a vector store and a relational database, with the documents imported properly. Inside the application you then need a hybrid procedure, partly deterministic and partly based on a language model: the question is analysed and, depending on the kind of request, handled in different ways.

The effort grows far faster than the precision you gain. The result is a complex processing flow, like the one below.

How a question flows

The language model steps in at three points at most; counts, sums, filters and citations remain deterministic code.

Codedeterministic stepModelcall to a language model

Preparation nothing is executed yet

  1. 01QuestionCode

    The natural-language question, with any explicit filters from whoever asks it.

  2. 02Value recognitionCode

    Looks in the question for values that actually exist in the data, such as names and codes, and for quantity mentions.

  3. 03Fast pathsCode

    The most frequent questions are recognised by deterministic rules, without calling a model.

  4. 04Coverage checkCode

    Is the whole question explained? If part of it is left uncovered, the fast path is discarded rather than giving a partial answer.

  5. 05ClassificationModel

    Only when no fast path is enough: the model says whether the question asks for an analysis or a search.

  6. Depending on the classification, one of two paths:

    06aCompilationCode

    For analyses, a deterministic compiler turns the question into a plan.

    06bPlanningModel

    For searches, the model proposes one to three steps as typed JSON. The code validates them: schema, the asker's filters, unused terms.

  7. 07Value resolutionCode

    Every value in the plan is checked against the real data. If it is unknown or ambiguous, a clarification is requested.

  8. 08Prepared planCode

    If the data changed in the meantime, preparation is repeated: an out-of-date plan is never executed.

Execution

  1. Depending on the kind of plan, one of two paths:

    09aAnalytics and aggregatesCode

    Counts, sums and analyses are computed by the database; text and citations are generated by code. No model involved.

    09bSearchCode

    Exact search in the database or semantic search on a vector index, then confirmed against the active data. At most 40 pieces of evidence.

  2. 10Answer draftingModel

    Searches only: the model writes the answer, citing the identifiers of the evidence it received.

  3. 11Citation checkCode

    The server checks every identifier and rebuilds the citations. An unknown reference makes the request fail.

  4. 12AnswerCode

    Text, status, executed plan and citations, with notes on the limits applied.

The other exits

  • Clarification or unsupported question. When the plan cannot be built with certainty: a deterministic message, never an invented answer.
  • Error. If validation fails, if the model cites evidence that doesn't exist, or if the data keeps changing: the request fails instead of returning a doubtful answer.

Where authority lies

  • The explicit filters of whoever asks the question override any filter proposed by a model.
  • Models return only schema-checked JSON; none of them writes SQL.
  • Counts, sums and analyses are computed by the database and presented by code: they never pass through the model that writes the answer.
  • Semantic search only finds candidates; the database confirms them only if they belong to the active data.
  • The answer may cite only evidence supplied by the server, and citations are rebuilt on the server side.

How it is developed

Every phase tested from the start

A flow like this has to be analysed and verified step by step from the very beginning. Otherwise you spend days writing code and, at the first test, the result is wrong with no idea where the problem begins.

Intermediate checks

Every step has its own check, based on data imported with a precise meaning: you always know at which point a result stops being correct.

The client's real questions

Development is driven by a set of realistic questions gathered with the client. In this scenario I don't think you can say "ask whatever you like and it will answer".

Online or local model

The method also matters because the result depends heavily on the model used, and neither option comes free.

Online services

The ChatGPT API or similar services: performance is better, but

  • you pay per call, and under heavy load the cost grows quickly;
  • privacy depends on what reaches the model. In this project the deterministic code extracts the sensitive part, and the model never sees it.

Open model on your own server

An open model on a GPU server avoids both problems, but

  • the hardware is expensive: over €5,000 for a 32 GB card;
  • with Qwen3.8 27B on 32 GB, it handles 2–3 concurrent requests;
  • a request takes 10 seconds or more, and execution has to be optimised;
  • smaller models lose precision, and many handle Italian poorly.

In my direct experience, all this calls for very precise analysis and development, with meticulous attention to the client's needs: the opposite of what you often read online.

My view on agentic approaches

I try to base my views on what I get my hands on and test. I have looked at and analysed agentic frameworks too, but given the infrastructure constraints described above, I believe that in ordinary settings it is very hard to propose a solution built on multiple calls to a model. It can make sense under three conditions:

  1. fast online services, at a relatively high cost;
  2. a higher cost for every request processed;
  3. use limited to high-value, and therefore infrequent, requests, not a tool open to any user, which would mean high spending for little return.

I don't think I'm the only one who sees this. Several companies, including large ones, are moving towards open models, hosted in the cloud or on data-centre hardware, combined with agentic architectures that have to be built and tuned to measure, because variability there is much higher.

In the end, "let's see what ChatGPT says" is no longer a strategy that holds up.

Considering AI in your applications?

Detailed CV and references available on request.