Lotus Tech Labs

AI and Machine Learning Development

Most AI features demo well and then fall over. The prototype answers in two seconds against a hundred documents, and then the real questions arrive: what happens at five million, what does a query cost, what does it return when the answer is not in the index, and who gets paged when the embedding job dies at 3am.

We build the feature and the system around it. That means the retrieval layer, the vector store and its index strategy, the evaluation that tells you whether a change helped, and the ordinary engineering underneath: queues, retries, caching, rate limits and a ceiling on spend. Python where the work is machine learning, Node and NestJS where it has to serve traffic, and the same people on both.

Retrieval that holds at scale
We run image similarity search over a library of roughly five million photographs, answering in under 200 milliseconds for about a million monthly active users. At that size the index strategy is the difference between a product and a timeout.
Answers that admit their limits
Returning nothing beats inventing something. We build the grounding, the citation path and the refusal case, because a confident wrong answer is the failure your users remember.
Evaluation, not opinions
A test set and a measured baseline before a prompt or model change ships, so "better" is a number you can compare rather than whoever argued hardest in the meeting.
Cost worked out per request
Inference and token cost modelled per request and capped before launch, so the bill does not become a surprise at the moment usage finally grows.

The process is the one we use for everything: a discovery call, a written PRD you approve before any code exists, agile iterations, then UAT sign-off. AI work differs in one way. The first iteration is usually an evaluation harness rather than a feature, because until there is one, nobody can honestly say whether the second iteration is an improvement. And if what you need is a model trained from scratch on your own research data, we are not the right team for it, and you will hear that on the first call rather than three weeks in.

Vector search and embeddings
Qdrant in production today, including the embedding pipeline, reindexing, and the backfill that nobody budgets for.
LLM features inside a running product
Added to a system that already has users, behind the same authentication, logging and rate limits as everything else in it, rather than bolted on beside it.
Automation where the work is repetitive
One marketplace we built for could not get vendors to write their own listings. A prompt now produces the description, the FAQs and the imagery, on n8n against their existing data.