The problem: the answer exists, but nobody can find it
A construction and surveying company runs on small facts. When the rebar arrives. What the last survey said. Which supplier quoted what, and when.
Those facts exist. They are just scattered:
- Voice notes recorded while walking the site
- Photos of a marked-up drawing
- Spreadsheets emailed once and never opened again
- PDFs from suppliers and the land registry
- Messages in a group chat that scrolls past by lunchtime
Ask "how much did we spend on rebar in September?" and the answer is in a spreadsheet. Ask "when is the next delivery?" and the answer is in a voice note from Tuesday. No one holds both.
So people ask a colleague. The colleague goes looking. Now two people have stopped working.
Why a search box doesn't fix it
Search finds documents, not answers. Type "rebar September" and you get five spreadsheets. You still have to open them and add the numbers up yourself.
And most AI document tools can't do sums. The common approach is to chop everything into passages, find the passages that look similar to your question, and hand them to a language model. That works for "what does our safety policy say". It falls apart on "how much did we spend".
A similarity search can find the rows. It cannot add them up. And it cannot tell you which row it used — so you cannot check it.
For a company that signs contracts based on these numbers, an answer you can't check isn't an answer.
What we built
An internal assistant, in Croatian, that staff use like a group chat.
Everything gets dropped into one place — typed messages, voice notes, spreadsheets, PDFs, photos. Anyone can then ask the chat a question and get an answer back, with a link to where it came from.

Two things happen to everything that arrives.
First, it is kept. Raw, immediately, before anything clever is attempted. Nothing is thrown away because a classifier had a bad day.
Second, a filter decides whether it is business information. "Delivery arrives Thursday" is. "Anyone want coffee" isn't. What passes gets processed twice over:
- As facts in a table — suppliers, dates, amounts, materials. This is what makes arithmetic possible.
- As searchable meaning — so "rebar" also finds "reinforcement steel", and a question phrased loosely still lands.
When someone asks a question, the assistant uses both. It runs a real database query for anything numeric, and a meaning-based search for anything descriptive, then answers from the two together.
Every answer shows its work
This was the part the client cared about most.
Under each answer there are source links. Not "according to your documents" — an actual link to the exact spreadsheet row, PDF page, or message the number came from. Tap it and you land on that row, highlighted.

It changes how people treat the tool. An answer you can check in one tap is one you'll act on.
Dates get pulled out of the conversation
Site work is scheduled in passing. "We'll do the survey on Thursday" is said once, in a voice note, and never written down anywhere else.
The assistant reads those mentions and puts them in a calendar, with a link back to the sentence it came from.

One detail mattered more than we expected. People routinely get half a date wrong — they'll say "on Sunday the 18th" when the 18th is a Friday. Rather than letting the AI silently pick one and sound confident, a mismatch is flagged for a human. A wrong date that quietly becomes a commitment is worse than no date at all.
The hard parts
Croatian is badly served by standard tools. Most text search engines ship with built-in support for English and a dozen other languages. Croatian isn't one of them. We had to hand-build the language handling — including a bug where đ survives the standard text-normalisation step untouched, so every search for a name like Đuro silently missed. There is now a test that exists purely to stop anyone "simplifying" that away.
The filter promotes, it never deletes. Original messages and files are always kept. The filter only decides what gets processed further. That means when we improve it, we can re-run it over everything already collected — which would be impossible if the first version had thrown things out.
It's a phone app, because the work happens outdoors. Site staff aren't at a desk. The app is built for one-handed use in a van or on a site: record a voice note, photograph a drawing, ask a question.

What this pattern is good for
This isn't specific to construction. The same shape fits any company where the knowledge is real but scattered, and where a wrong answer has consequences:
- Professional services — project history, scoped work, who agreed what
- Manufacturing — supplier terms, part specifications, incident history
- Property and facilities — maintenance records, costs per site, contractor history
- Any business running on a group chat — where the institutional memory is genuinely in the chat, and leaves when people do
The two things that make it work are worth repeating: combine real database queries with meaning-based search, and cite the exact source. Without the first, it can't count. Without the second, nobody trusts it.



