Featured work
Projects that best show the range of the work: retrieval and analysis, evidence-aware AI, deterministic decision support, operational workflows and end-to-end product building.
All 24 projects
Projects are grouped by the kind of problem they solve. Open any item for the complete original description, technical details and links.
Research, retrieval & evidence
02 Finnish guidance assistant that shows its sources Answers Finnish kevytyrittäjä questions from 14 official sources and shows the evidence behind the answer.
a RAG demo answering Finnish kevytyrittäjä (“light entrepreneur”) questions from 14 documents coming from Finnish authorities and pension and unemployment institutions. It shows supporting passages, publishers and update dates, excludes outdated guidance, and declines when evidence is insufficient.
Tech stack: TypeScript, Node.js, PostgreSQL with pgvector, OpenAI embeddings, Claude and Zod.
I compared semantic search with hybrid semantic-and-keyword search under identical conditions. Hybrid could be adopted only if it fixed a known failure without causing regressions. Twelve questions drove the decision; five were held back and run only afterwards.
On the original 11-document corpus, semantic scored 15/17 and hybrid 14/17. Hybrid fixed no known failure and introduced a wrongful refusal, so it failed the gate. In the held-back five, both methods failed another question: semantic declined, while hybrid gave the comparison's only confidently wrong answer.
After expanding the corpus to 14 documents, I reran the same questions and the ranking held. A different model graded the answers, and I checked every failure and version-sensitive answer (ten of 34) against the sources. I kept semantic search. The chatbot has additional safeguards not covered by this comparison.
03 Vector store RAG chatbot for questions related to a specific website Answers questions using information retrieved from a specified website.
(i.e. once you input the website name, the chatbot will give you information that can be found on the website). Tech stack: Langflow, Astra DB, OpenAI API.
22 Equity Research Snapshot Monitor Checks new company and market evidence against monitoring items in a dated equity-research report.
(TypeScript + Claude + Jev): built around a dated UPM equity-research report, the prototype checks new company and market evidence against analyst monitoring items such as Finnish wood costs, the WISA demerger and Nordic electricity prices. Deterministic code handles time boundaries, provenance and grounding; Claude identifies relevant evidence and what it still does not establish, without making investment or valuation judgments. I preregistered a small evaluation before the first live run and preserved four iterations as observed failure modes were fixed; the final Claude run accepted all 16 evidence/watch-item pairs and matched all 6 preregistered judgments. Once the main version was working reliably, I tested Jev on the first AI judgment: whether each document is relevant to a watch item and should be analysed further. Jev is built for this kind of classification rather than prose generation. It reproduced all 16 final Claude relevance decisions in just four requests, with each document checked against all four watch items at once. I kept the Claude-based version in the demo because it also provides the cited evidence and limitations an analyst needs, while the Jev test was too small to justify redesigning a working system.
AI generation, control & evaluation
01 Gateway that decides which messages an AI assistant is allowed to handle Routes requests before they reach an AI assistant, its documents or its model budget.
A gateway that decides which messages an AI assistant is allowed to handle. It lets a company put AI in front of customers or staff without every request automatically reaching its documents and its budget, and afterwards a receipt says what happened to any particular message and why. A list of literal patterns catches crude attempts to override the assistant's instructions before any model runs. Claude Haiku, far cheaper than the assistant it protects, judges the rest and permits, asks back, or refuses. The protected system here is a Finnish guidance assistant answering from a pinned document corpus, see project #2 below.
Tech stack: TypeScript, Node.js, Claude Haiku for routing, Zod, node:test.
Measured with and without the gate in front: legitimate requests come out slightly more expensive, not cheaper, because they pay for the decision and then for the work. What the gate removes is the cost of what should never have run: an override attempt cost nothing and reached nothing, where without the gate it took 7.9 seconds and $0.045. Whether that nets out depends on how much of your traffic should be blocked, a fact about your traffic, so no savings figure appears anywhere.
The router is measured against ten labels written before the first run; it currently matches all ten, which is a small sample and not a guarantee.
It does not claim to stop prompt injection, and a receipt records a wrong decision as faithfully as a right one.
17 EU-Sovereign AI EIC Accelerator Short Proposal Drafter Turns founder inputs into a criteria-mapped EIC Short Proposal draft on EU infrastructure.
(Next.js + Mistral AI, EU-hosted): an AI proposal assistant built around one constraint: an EU company should never have to send confidential proposal data through US-hosted AI services. It turns a founder's inputs (voice or text, plus a pitch deck) into a criteria-mapped first draft of every official EIC Short Proposal section (Excellence, Impact, Risk, Video Script, Pitch Deck) in roughly 3 to 5 minutes. On top of the draft sits an evaluator that scores each section against the real EIC requirements (covered, partial, or missing), flags the weak spots, and lists the critical fixes before submission, with a realism check that keeps the draft grounded in the applicant's actual data instead of inventing evidence. The whole pipeline stays on EU infrastructure end to end: EU-hosted Mistral on the user's own key (data goes straight from the browser to the model, nothing pooled or retained in between), voice recordings auto-deleted after 5 minutes, pitch-deck parsing and PDF generation done in-app with no external processor, and downloads behind authenticated, time-limited links. Stack: Next.js (Vercel EU region), Supabase EU West region (row-level security, audit logging, signed URLs), Mistral AI, Mistral Voxtral voice transcription, Apify (EU) for website analysis, Resend for email.
20 Meeting-to-Proposal Generator, voice-matched Turns a discovery-call transcript into either a proposal or a scoping brief in the target agency's voice.
— the hand-built counterpart to #16 (custom Python, pure-stdlib SSE server, Mistral EU-sovereign, no framework, no SaaS glue): discovery-call transcript → two LLM gate decisions (qualify the lead + priced proposal vs. no-commitment scoping brief) → proposal streamed live in the target agency's own voice, learned from a public writing-style profile rather than slot-filled into a template → voice-match panel pairing real public quotes with the generated phrasing. Hardened against AI slop: no fabricated cases/numbers, scoping-brief path forbids fixed prices, Art. 50 transparency footer, human-verify-before-send. Synthetic transcript, no PII, EU-sovereign model (no US vendor). Where #16 orchestrates existing tools to fill a template fast, this one is built from scratch for sovereignty, on-voice generation, and judgment (gates) over slot-filling.
21 Multilingual Catalog Enrichment with a Safety Check Generates listings in five languages while independently blocking unsupported technical claims.
(Python + Claude): messy supplier product row → Claude rewrites it into clean listings in five languages (Finnish, English, Estonian, Latvian, Lithuanian) → a separate checking step reviews the result before it ships. A 13-product proof of concept for an outdoor-gear catalog, where the real risk is not clumsy copy but invented safety numbers: asked about a ski, the model claims "waterproof up to 15000 mm", the checking step catches that no supplier ever provided that figure, removes it, leaves the field blank, and flags the product for a person, while a rain jacket's genuine 20,000 mm rating (which the supplier did provide) is kept and marked as verified. Safety and technical numbers only ever come from supplier data, never from the model, and that rule holds even when the model ignores its instructions, because a second layer enforces it independently. Every field is labelled by where it came from; anything missing goes to a human review queue instead of being guessed. Cost measured on the real run at about one cent per product, so roughly €79 projected to enrich all 8,000 once. Live browser demo you can watch on a single product. Stack: Claude (Anthropic API), plain Python, static web host.
Operational data & decision support
18 Data Reconciliation Engine Reconciles conflicting CRM and operational data using explicit field-level source rules.
(Python, deterministic, offline, no LLM): two disagreeing systems (CRM + ops sheets) → entity resolution (normalization + alias map + fuzzy match) → rule-based conflict resolution with explicit system-of-record per field — e.g. contract value resolves to the CRM, delivered impact resolves to the ops sheet, not a blanket "trust one system" — and full provenance → staleness/single-source flags → one number auto-derived into 4 external report formats (scorecard, investor, EcoVadis-style ESG, CBAM-style). Turns conflicting source systems into one answer-ready portfolio for due-diligence/ESG/regulatory reporting; conflicts resolved by written rules, not a model; humans verify flagged items. (Synthetic data; transformation pattern.)
19 Voice-controlled agentic layer Lets a voice agent query reconciled business data without allowing the model to invent or resolve figures.
— a multi-turn voice assistant built on top of #18 (Python, provider-agnostic function calling, self-hosted Whisper STT + Piper TTS, Mistral EU).
Ask a question in Finnish → local Whisper transcribes it → the model chooses which read-only tool to call → local Piper speaks the answer. Unlike a fixed workflow, the model selects its tools based on the question and can choose a different route as the conversation develops.
Ask why a client's impact figure resolved to delivered tonnes rather than the sales pledge, then follow up about contract value. The agent retrieves the relevant figures and rules from #18 and explains both decisions.
The boundary: the model chooses the tools, but deterministic code owns every number and reconciliation rule. The agent never invents figures or decides which source wins.
23 Material Certificate Reconciler Matches received material against supplier certificates and sends ambiguous or conflicting cases to a person.
a deterministic demo on fictional data. Material keeps arriving, and so do supplier certificates. The system checks whether the evidence actually identifies and covers the received material, resolves only the cases where exactly one certificate establishes the relationship under an explicit rule, and sends the missing, conflicting and ambiguous remainder to a person rather than guessing.
Stack: Python, Flask/Jinja2, pypdf, reportlab; pytest for acceptance testing. No database or frontend build step.
24 PO Confirmation Reconciler Checks supplier confirmation PDFs against ERP purchase orders and surfaces the exceptions that need a buyer's attention.
PO Confirmation Reconciler: the same underlying idea as in the Material Certificate Reconciler, applied to a different operational process: a deterministic demo on fictional data where supplier confirmation PDFs are checked against open purchase orders from an ERP export. The system surfaces what changed or is missing, uses buyer-set tolerances to decide which date-related differences need attention and how urgently, and separates cases that cannot safely be compared, such as unit or currency mismatches. When supplier documents contradict each other, both claims are kept as evidence and the buyer decides. Together, the two builds apply the same evidence → exception → human judgment pattern across different operational contexts and decision rules.
Sales, marketing & business-process automation
04 Quoting Process Prototype Generates quote drafts from Airtable data using predefined templates.
(n8n): Airtable → n8n → OpenAI → Google Docs; generates quote drafts using predefined templates (results in about 30% time reduction).
05 CV Evaluation Prototype Extracts CV content, scores applications with AI and writes the results to Google Sheets.
(n8n): Form data → n8n → PDF text extraction → OpenAI scoring → Google Sheets; automates initial application screening.
06 CRM Follow-up Triggers automated customer follow-up from changes in a ClickUp CRM workflow.
(Make.com): ClickUp status updates → Make.com → automated customer follow-up messages; improves lead tracking and engagement.
07 Web Scraping Workflow Collects listing data overnight and writes the results to Google Sheets.
: ScraperAPI + Make.com → collects listing data to Google Sheets – overnight market data updates. Adaptable e.g. for e-commerce site monitoring.
08 Lead Intelligence Feedback Loop Adds AI-generated sales insights and icebreakers to new Salesforce leads.
(Make.com): Salesforce Lead creation → Make.com → Gemini API → Salesforce record update → Google Sheets – automatically generates within 5 s AI-powered sales insights and icebreaker suggestions for new leads. It eliminates the research scramble, increases the amount of first-reply rates and ensures no lead slips through.
09 Automated Email Autoresponder System Analyzes incoming lead emails and generates personalized CRM-connected responses.
(Make.com): New lead email trigger → Make.com → OpenAI GPT-4 content analysis → 79-90 second human-like delay → AI-generated personalized response → CRM integration (ClickUp) – automatically delivers professional, customized replies to prospects within minutes, improving conversion rates by 5-10% and enhancing customer experience for coaching companies and agencies.
10 Photography Client Management System Automates booking administration, customer emails, billing and reminders.
(Make.com): WordPress booking webhook (mocked) → Make.com → ClickUp CRM task creation → automated email sequence (confirmation, bill, preparation list) → daily scheduled reminders (7-day email + 1-day SMS via Twilio) – streamlines photography business operations by automating client communication, booking management, and follow-up reminders to enhance customer experience and reduce manual administrative work.
11 Automated Proposal Generation System Generates a proposal from CRM data by filling a Google Slides template automatically.
(Make.com): Monday.com stage change webhook ("generate proposal") → Make.com → Monday.com data query → Google Slides template auto-fill → shareable link generation → Monday.com record update (proposal link + stage to "Discovery") – eliminates manual proposal creation, saving sales teams 20-30% of their time and companies 5-10% margin by automating document generation and allowing focus on high-leverage selling activities.
12 Sales Rep Performance Analyzer Turns weekly sales data into AI-generated coaching tasks inside Salesforce.
(Make.com): BigQuery weekly sales data → Make.com → Salesforce contact verification → Gemini AI coaching insights → Salesforce task creation → BigQuery audit logging – automatically generates personalized AI coaching in < 30 s for sales reps based on their weekly performance (deals closed, revenue), delivering actionable insights directly as CRM tasks while maintaining complete audit trails for compliance and performance tracking.
13 Campaign Copy Generator Combines Salesforce campaign data with BigQuery segments to generate campaign copy automatically.
(Make.com): Salesforce new campaign trigger → BigQuery segment data query → Gemini AI copy generation → Text formatting & validation → Salesforce campaign update → BigQuery audit logging. It automatically generates personalized marketing copy in ≈20 s (vs ~2 h of manual writing) for new campaigns by combining real-time customer segment analytics (size & purchase recency) with AI creativity, delivering ready-to-use email subject lines and body content directly into CRM campaigns while maintaining complete audit trails for content governance and performance analysis.
14 Lead Scoring & Assignment Automation Scores new leads from behavioral data and writes the recommendation and next action back to Salesforce.
(Make.com): Salesforce new-lead trigger → BigQuery behavior query → Gemini AI scoring & recommendation → score-based router → Salesforce lead update + task creation → BigQuery audit logging. The flow qualifies each lead in ≈10 s (vs 30 min manually), writes a transparent score, next action and task, and stores a full audit trail, while ensuring hot leads get instant attention and cold ones move to nurture.
15 Sales Agent Follow-Up System Runs AI-personalized email follow-up sequences until a lead books a call or requires re-engagement.
(n8n): Google Sheets Lead List → Schedule Trigger → n8n → Google Sheets CRM → OpenRouter AI → Gmail → cal.com integration – fully automates lead management by transferring contacts from Lead List to CRM, executing AI-personalized 3-email follow-up sequences every 2 days until call booking, automatically updating CRM when prospects schedule meetings, and re-engaging no-show leads with targeted follow-ups. The system ensures zero leads slip through, and delivers contextual AI-written emails that improve response rates by 15-20% for agencies and consultants.
16 AI Proposal Generation Turns a sales-call transcript into a formatted proposal using an existing Google and CRM stack.
via no-code orchestration (n8n): Gmail trigger (Google Meet transcript) → n8n → Google Docs transcript extraction → Google Sheets CRM lookup → OpenRouter AI (Claude) information extraction → Google Slides template duplication → automated text replacement → Gmail review notification → CRM update – wires a team's existing SaaS tools into one pipeline that turns a sales-call transcript into a fully-formatted, template-based proposal in ≈45-60 s by extracting 18+ data points (bottlenecks, solutions, phases, pricing), filling a Google Slides template, and routing the draft for human review. Eliminates 2-3 h of manual proposal writing per deal, zero information loss from calls. Strength: stands up fast on top of an existing Google/CRM stack, no custom code.