Every phase tested from the start
A flow like this has to be analysed and verified step by step from the very beginning. Otherwise you spend days writing code and, at the first test, the result is wrong with no idea where the problem begins.
The most complex case I have worked on so far is a hybrid RAG system: an application that, starting from imported delivery documents, lets you query business data in natural language. It is the classic "chat with your data", with one requirement that changes everything: precision.
If you ask which products sold best over a period, the result must match what the management system says.
A classic vector store, with semantic search alone, cannot guarantee that precision. You need both a vector store and a relational database, with the documents imported properly. Inside the application you then need a hybrid procedure, partly deterministic and partly based on a language model: the question is analysed and, depending on the kind of request, handled in different ways.
The effort grows far faster than the precision you gain. The result is a complex processing flow, like the one below.
The language model steps in at three points at most; counts, sums, filters and citations remain deterministic code.
Codedeterministic stepModelcall to a language model
The natural-language question, with any explicit filters from whoever asks it.
Looks in the question for values that actually exist in the data, such as names and codes, and for quantity mentions.
The most frequent questions are recognised by deterministic rules, without calling a model.
Is the whole question explained? If part of it is left uncovered, the fast path is discarded rather than giving a partial answer.
Only when no fast path is enough: the model says whether the question asks for an analysis or a search.
Depending on the classification, one of two paths:
For analyses, a deterministic compiler turns the question into a plan.
For searches, the model proposes one to three steps as typed JSON. The code validates them: schema, the asker's filters, unused terms.
Every value in the plan is checked against the real data. If it is unknown or ambiguous, a clarification is requested.
If the data changed in the meantime, preparation is repeated: an out-of-date plan is never executed.
Depending on the kind of plan, one of two paths:
Counts, sums and analyses are computed by the database; text and citations are generated by code. No model involved.
Exact search in the database or semantic search on a vector index, then confirmed against the active data. At most 40 pieces of evidence.
Searches only: the model writes the answer, citing the identifiers of the evidence it received.
The server checks every identifier and rebuilds the citations. An unknown reference makes the request fail.
Text, status, executed plan and citations, with notes on the limits applied.
A flow like this has to be analysed and verified step by step from the very beginning. Otherwise you spend days writing code and, at the first test, the result is wrong with no idea where the problem begins.
Every step has its own check, based on data imported with a precise meaning: you always know at which point a result stops being correct.
Development is driven by a set of realistic questions gathered with the client. In this scenario I don't think you can say "ask whatever you like and it will answer".
The method also matters because the result depends heavily on the model used, and neither option comes free.
The ChatGPT API or similar services: performance is better, but
An open model on a GPU server avoids both problems, but
In my direct experience, all this calls for very precise analysis and development, with meticulous attention to the client's needs: the opposite of what you often read online.
I try to base my views on what I get my hands on and test. I have looked at and analysed agentic frameworks too, but given the infrastructure constraints described above, I believe that in ordinary settings it is very hard to propose a solution built on multiple calls to a model. It can make sense under three conditions:
I don't think I'm the only one who sees this. Several companies, including large ones, are moving towards open models, hosted in the cloud or on data-centre hardware, combined with agentic architectures that have to be built and tuned to measure, because variability there is much higher.
In the end, "let's see what ChatGPT says" is no longer a strategy that holds up.