Neuraflow / Applied AI engineering

Building the RAGfoundation behindmunicipal assistants.

I worked with Neuraflow's CTO on the ingestion, retrieval, deployment and evaluation infrastructure behind public-facing assistants used across approximately 20 German municipal customers.

Generative AI Engineer · April 2024 — March 2025

municipal customers
~20
data refreshes
Weekly
multi-tenant retrieval
Isolated
production deployments
Public

Public deployment · Bonn

City of Bonn website open on a laptop with its red municipal AI assistant answering an identity-card question
A Neuraflow assistant embedded directly into the City of Bonn website. The city has published an article about the deployment.

01Context

Municipal information is difficult to navigate.

Residents should not need to understand a municipality’s information architecture before they can ask, “How can I register my dog?” or “My passport is expiring. What do I need to do?”

The assistants let people ask those questions in normal conversational language, across multiple languages. Answers were grounded in each municipality’s own service information and delivered inside its public website.

02System

A shared platform, isolated municipal data.

We did not build an independent RAG stack for every customer. One reusable platform supported separate municipal deployments, while each municipality’s retrieval data remained isolated.

Source formats varied by customer. The ingestion layer turned structured APIs, scraped pages and documents into a common path toward retrieval.

Multiple municipalities feed a shared ingestion and retrieval platform, while their datasets and resident-facing assistants remain separate.

Conceptual view; not a literal database topology

03Architecture

Retrieval quality started at the source.

The LLM was one stage in a longer production path. Most of my work lived in the layers that determined what evidence reached it.

Structured APIs
Often XML-based municipal service data
Municipal websites
Primarily collected through Apify scraping
Documents
PDFs and other files entering the same retrieval path

Municipal APIs, websites and documents pass through ingestion, normalization, chunking, OpenAI embeddings, Pinecone indexing, retrieval, prompt construction and an OpenAI language model before a resident receives an answer.

Ingest
APIs · Apify · documents
Normalize
clean into a common shape
Chunk
smaller retrieval units
Embed
OpenAI embeddings
Index
Pinecone · isolated data
Retrieve
query-relevant context
Generate
prompt · OpenAI LLM
Answer
resident-facing response

Weekly refresh jobs · Azure · Docker · Terraform · Monitoring · Development environment

04Ingestion

The model wasn't the hardest part.

The most common input was structured municipal data exposed through APIs, often as XML. Other customers depended on website content, PDFs and general documents. Apify handled most website collection.

These sources did not arrive in one clean, consistent format. The engineering work was turning different inputs into dependable retrieval material without assuming every municipality published information the same way.

05Retrieval

A correct page can still be bad context.

Large municipal service pages often mixed unrelated topics. Vector search could retrieve the broadly correct page for “How do I register my dog?” while burying the useful passage inside a large amount of irrelevant context.

We reduced oversized page-level documents into smaller retrieval units. A large page could become roughly four or five pieces, making returned context more focused without relying on the model to find one useful detail in a wall of text.

Before · one retrieval unit
Waste collectionDog registrationParking permitsIdentity cardsEvent permits

Correct page, noisy context.

After · smaller retrieval units
WasteDog registrationParkingIdentityEvents

Focused evidence for the answer.

The model could not compensate for consistently poor retrieval.

One of the largest improvements came from changing how information entered the vector database, not from changing the LLM.

06Freshness

Public information does not stay static.

Opening hours, administrative processes and municipal pages change. Scheduled jobs refreshed source data approximately once per week; this was a recurring ingestion process, not a one-time knowledge import or real-time synchronization.

Refresh cadence ~7 days

07Evaluation

Debugging RAG means finding the failing layer.

Evaluation was primarily manual: run a repeatable set of questions, inspect retrieved content, check the answer and use production feedback. A bad answer was not automatically an LLM failure.

  1. Source
  2. Ingestion
  3. Chunking
  4. Retrieval
  5. Prompt
  6. Generation

08Production

Public-facing, not a prototype.

Residents used these assistants inside real municipal websites. Delivery therefore included more than model and retrieval code: repeatable Azure infrastructure, Docker, Terraform, scheduled jobs and monitoring.

A separate development environment let the team test prompt and behaviour changes before collaboratively deploying updates to public systems.

Public deployment · Detmold

City of Detmold website with the green ChaDT assistant open, showing multilingual messaging and suggested municipal-service questions
A second deployment on the City of Detmold website, showing the shared platform adapted to another municipality.

09Role

What I owned

I worked closely with Neuraflow’s CTO throughout the architecture and delivery of the platform. My responsibility was substantial and collaborative — not sole ownership of the company’s entire system.

Primary responsibility

  • Ingestion architecture
  • Retrieval
  • Prompt construction
  • Deployment
  • Monitoring
  • Scheduled refresh jobs
  • Evaluation

Contributed to

  • Vector and data schema
  • Backend and API development
  • Internal admin tooling

10Takeaway

The pipeline was the product.

Production RAG changed how I thought about applied AI. Answer quality depended on source ingestion, document structure, chunking, tenant isolation, retrieval, prompt behaviour, freshness, evaluation and deployment infrastructure.

The biggest practical lesson was that improving a RAG system often meant fixing the pipeline before changing the model.

Stack

AI
OpenAI LLMs · OpenAI embeddings · RAG · Pinecone
Data
Python · XML APIs · Apify · web scraping · document processing
Infrastructure
Azure · Docker · Terraform · scheduled jobs
Engineering
Retrieval · prompt engineering · monitoring · evaluation · APIs