← All projectsCase study · 01

Newmann

An AI email assistant that reads your inbox the way you do: it drafts replies in your own voice, and lets you automate labels and rules either by building them or by describing them in a chat.

Role

Founding Engineer & Tech Lead

Type

Newmann AI

Year

2024 to Present

Category

Web Development · AI Engineering

Next.jsReactTypeScriptJava Spring BootOpenAIPineconeSupabasePostgreSQLGmail APIMicrosoft GraphOAuth2AzurePostHogResendGitHub ActionsVercel
Newmann inbox dashboard shown on a laptop mockup
Overview

Newmann is a B2B SaaS platform that connects to Gmail or Outlook and turns an inbox into something manageable: it decides what actually deserves attention, writes draft replies grounded in how you have answered similar emails before, and applies the labels and rules you define, either in a rule builder or by describing them to a chatbot. The platform is currently in testing ahead of its first release.

My role

I joined at inception as Founding Engineer and Tech Lead. I own the architecture and the technology choices, from stack selection to the design of the REST API in Java Spring Boot, and the AI engineering end to end: every prompt is written from scratch and the retrieval pipeline on Pinecone is mine. I work alongside a front-end and a back-end developer who carry the implementation forward, and the decisions that shape the system come through me.

The challenge

A generic model can write a polite email. It cannot write your email. The product only works if the draft sounds like the person sending it and if the assistant knows when to stay silent, and both of those have to hold while processing every message that lands in an inbox, at a cost per email that a subscription can absorb.

01

Drafts had to be grounded in each user's own history, not in a generic model voice

02

Gmail push notifications can arrive more than once, so the pipeline had to be idempotent and never produce duplicate drafts, labels or vectors

03

Newsletters and automated senders are a large share of any inbox: running an LLM on all of them would have burned budget for nothing

04

Two providers, Gmail and Microsoft, each with its own OAuth flow and its own idea of what a label is

05

Users needed both precise control over automations and a way to create them without learning a rule syntax

Process

How it came together

Newmann settings: theme, language, email signature, Newmann labels and the user's role
Editing a draft in Newmann: the original draft next to the new version, with a field to ask the AI to regenerate it
Newmann automation chat creating a label from a natural language description
The labels list in Newmann, with the automated drafts setting and the rules linked to each label
01

Architecture and the ingestion pipeline

Spring Boot exposes the REST API, PostgreSQL on Supabase stores the data with row-level security and multilingual content, and Next.js with React and TypeScript runs the front-end. Incoming mail arrives through a webhook pipeline built around idempotency and deduplication keys, so the same message can be delivered twice without ever producing a second draft. Providers are modelled per email account rather than per provider type, which is what made adding Microsoft alongside Gmail a configuration change instead of a rewrite.

02

Retrieval, prompts and the user's voice

Every relevant email is embedded and stored in Pinecone, so a new message is answered with the user's own past exchanges in context. I keep separate namespaces for context emails and for feedback signals: when the two shared a namespace, rejected drafts started polluting retrieval and pulling the model towards the answers the user had explicitly turned down. Rejections are also categorised, because 'this needed no reply' and 'the tone was wrong' are two different lessons and only one of them should stop future drafts.

03

Two ways to build an automation

The rule builder gives full control: you define the label, the conditions and the written description the model works from, and you can see exactly what will happen. Next to it sits a chatbot for everything else: you describe what you want and a stateful multi-turn flow assembles the same rule, asking for the missing pieces and stopping to confirm when it looks like one that already exists. Both paths write to one rule model, so nothing behaves differently depending on where it was created.

04

Shipping it and keeping it observable

CI runs on GitHub Actions for both front-end and back-end, with branch protection; the front-end deploys to Vercel and the Spring Boot API runs on Azure, with Flyway migrations against Supabase. I set up product analytics on PostHog EU and transactional email on Resend; the deployment pipeline was built with another developer and the team maintains it together. Getting the deployed environment stable meant working through OAuth redirects, CORS, HikariCP pool sizing and Supabase connection routing over the session pooler, the unglamorous half of running your own infrastructure.

Key decisions

Choices that shaped the product

Detect automated senders from headers first, AI only as a fallback

Why

Newsletters and no-reply senders identify themselves in the email headers. Reading them costs nothing and covers around 95% of cases, so the model is only called for the genuinely ambiguous ones. Those emails are still embedded into Pinecone as context, and only draft generation is skipped.

Trade-off

A hand-written heuristic to maintain as senders change how they label themselves.

Separate Pinecone namespaces per retrieval purpose

Why

Context and feedback answer different questions. Keeping them apart is the difference between a draft informed by how you write and a draft drifting towards what you already rejected.

Trade-off

More namespaces to manage, and every new signal type needs a deliberate decision about where it lives.

A rule builder and a chatbot, not one or the other

Why

People who know exactly what they want should not have to negotiate with a chat, and people who do not should not have to learn a form. Both surfaces produce the same rule, so the choice is about comfort rather than capability.

Trade-off

Two interfaces over one model: every change to what a rule can do has to land in both, and stateful multi-turn conversations are far harder to test than a form.

Parallelisation and caching over a heavier AI framework

Why

The response time users feel comes from how many model calls run at once and how many are avoided entirely. CompletableFuture parallelisation, batching and a lazy cache for importance evaluation moved the numbers; an extra abstraction layer would not have.

Trade-off

More concurrency to reason about, and caching means being explicit about when a stale verdict is acceptable.

EU-hosted analytics and infrastructure from the start

Why

Email content passes through the system and the customers are European. Choosing EU-hosted services while the codebase was small made data residency a setting rather than a migration.

Trade-off

A narrower set of providers to choose from, sometimes at a higher price.

Results

What changed

~95%

of automated senders identified from email headers alone in testing, at zero token cost

Gmail · Outlook

Both providers supported through a per-account model, each with its own OAuth2 flow

Idempotent

A webhook pipeline keyed so that repeated deliveries cannot produce a second draft, label or vector

2 paths

A rule builder and a conversational assistant writing to one automation model

Gallery
Newmann sign-in page, with Google or Microsoft login
Newmann dashboard with labels, automated drafts and active rules
Newmann landing page
What I learned

Deciding the architecture first means every shortcut becomes somebody else's inheritance. Isolating providers per account, keeping retrieval namespaces separate and making the pipeline idempotent all looked like over-engineering on day one, and they are the reason two more developers could build on top without renegotiating the foundations.

The cheapest AI call is the one you do not make. Reading a header before reaching for a model changed the unit economics of the product more than any prompt optimisation did.

Feedback is not one signal. Treating 'no reply needed' and 'wrong tone' as the same rejection quietly taught the system to stop being useful, and separating them was a modelling problem long before it was a prompting one.

Offering two ways into the same feature was worth the duplicated surface, but only because both write to a single rule model. Had I let the chatbot build its own shortcut version, the two would have drifted within a month.