Work
EN FR
Case Study · loicsans.me (independent project) · 2026 · 10 min read

It reads every word of this site, and it still won't tell you to hire me.

A conversational assistant grounded in this site's case studies and journal, in English and French. There is no retrieval step, the whole corpus goes into context on every question. It says when it doesn't know, and it refuses the one question it would be most rewarded for answering. Live at loicsans.me/ask.

Ask it something
Client
loicsans.me (independent project)
Role
Designer & Builder (Solo)
Year
2026
Disciplines
AI ProductConversational UXPrivacy by DesignFront-end EngineeringR&D0→1 Product
It reads every word of this site, and it still won't tell you to hire me.

The question a portfolio never answers

A recruiter with six minutes and eleven tabs open has one question, and it is never "walk me through your process". It is "has he done fintech", or "can he ship or does he only draw", or "is he free in March". A portfolio makes them read six case studies to find out. Most of them don't. They skim two, guess, and close the tab. So I built the thing that answers directly: an assistant that has read every word of this site and will talk about it, in either language. Then I spent most of the project deciding what it must refuse to do.

The constraint nobody imposed

No client, no budget holder, no governance review. That is the trap, not the freedom. A portfolio assistant with nobody to answer to becomes a salesman, and a salesman on your own site is worth less than nothing, because every visitor knows who wrote the prompt. So I set the constraint myself and made it structural rather than aspirational. The assistant may describe what the site says. It may say it doesn't know. It may not tell you what to conclude about me.

Grounding, not persuasion

The failure mode here is not being wrong. It is being fluent, plausible and unverifiable about someone's career, in the one place where the visitor has no way to check. So the design problem was never "answer questions well". It was how to build something that stays tethered to a source, admits the edge of what it knows, and declines the question it has every incentive to answer.

The AI-enhanced pipeline

Four stages. One principle: the model's output is untrusted input, all the way through.

STAGE 01

Deciding what it refuses

The problem
An assistant on a portfolio has an obvious incentive: sell. Ask any competent model whether its author suits your role and it will produce a fluent yes. It will be persuasive, and it will be worthless, because it is my site and my instructions.
The method
Wrote the refusals before the capabilities. Three of them. It will not predict how I would perform somewhere it knows nothing about. It will not present its own inference as a claim I made. It will not fill a gap with a plausible guess, it names the gap and points at my email. The system prompt was then built around those, with a smoke suite that probes each one against the live endpoint.
AI handles
Drafting candidate prompt language, and generating adversarial questions designed to bait a sales answer out of it.
I own
The decision that refusing to sell is the feature, not a limitation to apologise for. A portfolio that tells you what to think about its author has told you nothing. The credibility comes from the assistant being willing to say the site does not cover that.
The result
Ask it whether I am right for your role and it declines, describes what I have actually done, and gives you my email. That is an answer a hiring manager can use.
STAGE 02

Whole site in, no retrieval

The problem
The reflex for "chat with my documents" is retrieval: chunk the corpus, embed it, pull the three most similar passages, answer from those. It is the default because it is what you do when the corpus is too big for the context window. Nobody checks whether theirs is.
The method
Measured first. Every case study, journal post and page, plus a hand-maintained facts file, in both languages, came to 66,443 tokens. That fits comfortably. So the entire site goes into the system prompt on every question, cached between turns so the second question costs a fraction of the first.
AI handles
Building the corpus extractor across three different bilingual conventions in the content files, and the token counting harness.
I own
The judgment that retrieval was the wrong tool at this size. Retrieval answers from the passages it chose, which is how a confident answer gets built on the wrong page. At 66k tokens the retrieval step buys a cost saving and costs accuracy. Below some threshold the sophisticated architecture is the worse one.
The result
No chunk is ever missed, because there is no chunking. A question spanning four case studies is answered as easily as one about a single project.
STAGE 03

Treating the output as untrusted

The problem
The model returns markdown. Rendering it means turning text a language model produced into HTML in a visitor's browser, which is an injection surface whether or not you think of it that way.
The method
Wrote a tokenizer that recognises exactly two things, internal links and bold, and escapes everything else. Then wrote the tests that try to get past it before trusting a line of it.
AI handles
The first draft of the tokenizer and the initial test cases.
I own
Deciding the allowlist was two constructs rather than a markdown library. A parser that supports images, raw HTML and arbitrary protocols is a larger surface than this feature needs, and every construct I did not support is a class of bug I did not have to reason about.
The result
Answers render in the site's own typography instead of raw asterisks, and the render path carries a test that specifically tries to break out of it.
Honest note
The first version of the link pattern allowed protocol-relative URLs. A model talked into emitting a link beginning with two slashes would have produced a working off-site link on my own domain. The test caught it, not me reading the code.
STAGE 04

Cheap enough to leave running

The problem
Putting the whole site in context on every question is the right architecture and the expensive one. The first version ran on the largest available model at roughly $0.54 a conversation. An afternoon of testing cost $5. At that price you quietly switch it off.
The method
Two changes, measured rather than guessed. Moved to a mid-tier model after confirming answer quality held on the probe set. And found a real defect: the corpus splitter only handled bilingual keys at the top level, so nested French text was leaking into the English corpus and back. Fixing it took the corpus from 86,762 tokens to 66,443.
AI handles
Running the cost arithmetic across model tiers and cache read and write rates, and building the token counting probe.
I own
The decision to measure before optimising. The exact count came from the API's own counting endpoint, which is free and precise. The estimate I had been reading off the build log was running 25 percent low, which is the sort of error that sends you optimising the wrong thing.
The result
$0.25 a conversation, down from $0.54, and 77,284 tokens today. A hard ceiling of 20 conversations a day across the whole site caps a bad day at roughly $6.
Honest note
Then writing this case study put it back up. The corpus is in the corpus, so these 8,104 tokens are billed on every question anyone asks, and a cold conversation now costs $0.29 rather than $0.25. Teaching the builder to read the /ask page, which the assistant had never been able to see, added another 979. Both are worth paying for. Neither was free, and that is exactly the kind of drift a system stops reporting the moment you stop measuring it.

The address that did not exist

loicsans.me · Site Assistant · 2026

Testing it one afternoon, I asked how to get in touch. It gave me an email address at loicsans.me. That address has never existed. My email is at loicsans.com, and nothing in the corpus said otherwise. The model had seen the domain it was running on, seen a contact section, and produced the address that should have been true. This is the failure that matters in a grounded system, and it is not the one people design for. It is not a hallucinated fact you catch by reading, because it is not wrong in any way a reader would notice. It is right-shaped. A visitor emails it, the message bounces, and the last thing my portfolio did was waste their time. The fix took two lines: pin the address in the prompt, and add a probe that asks for it on every smoke run. The lesson took longer. Grounding a model in a corpus does not stop it filling gaps, it only narrows where the gaps are. Every fact a system cannot afford to have invented has to be stated rather than inferred, and then tested on every run.

Where AI ends and judgment begins

Three places the model was no help, and one where it was actively misleading. It could not tell me what to refuse. Asked to draft the refusals, it produced hedging language that sounded careful and committed to nothing. It could not judge its own answers. I wrote the smoke suite three times. The first two versions asserted on phrasing and failed correct refusals for using different words. A test that checks vocabulary instead of behaviour is worse than no test, because it fails the good runs and passes the bad ones. The final version asserts the property, and I validated the predicate offline against three real refusals and three fabricated answers before spending another API call on it. And it was confidently wrong about its own contact address, which is the story above.

What it delivered

Live, and it says no

It runs at loicsans.me/ask in both languages. Ask whether I am right for your role and it declines, describes what I have actually done, and hands you my email. That refusal is the part I would defend hardest.

Cost halved, with a bug behind it

$0.54 a conversation down to $0.25, through a model change and a corpus splitter that had been silently duplicating French text into the English corpus. 86,762 tokens became 66,443. The saving was real, and half of it was a defect rather than an optimisation. Adding this case study has since taken it to 77,284, because the corpus contains the corpus.

A privacy promise the code enforces

The question and the date, nothing else. No address, no hash, no session id, no way to group one person's questions together. Deleted at 90 days by a cleanup whose key pattern is written strictly enough that it cannot reach the rate-limit counters. The /ask page says all of this in plain language rather than burying it.

56 tests, including the ones that hurt

The render path is tested against a link that tries to leave the domain. The prompt is probed for the address it once invented. The corpus splitter has a test for the nesting bug that shipped.

Decisions log

The calls that mattered, and what changed my mind. The reversals stay in the record, they are the part with something to teach.

01
Put the entire site in context on every question

Instead of: Retrieval over an embedded, chunked corpus

At 66k tokens the whole corpus fits. Retrieval would answer from the three passages it happened to pull, which is exactly how confident answers get built on the wrong page.

02
Refuse to assess my own fit for a role

Instead of: Answer it persuasively, it is a portfolio after all

An assistant that recommends its own author has told the visitor nothing they can trust. The refusal is what makes everything else credible.

03
Record the question and the date, nothing else

Instead of: A session id, so conversations could be read as threads

I wanted to know what the site fails to explain. That needs the question, not the person. A thread id would have made every line traceable in exchange for a benefit I did not need.

04
Move from the frontier model to the mid-tier one Reversed

Instead of: Stay on the largest model for answer quality

$0.54 a conversation is a switch you end up turning off. Quality held on the probe set, so the premium was buying nothing I could measure.

05
Cap the whole site at 20 conversations a day

Instead of: A generous per-visitor limit and nothing global

Every cold conversation pays to cache the corpus. A per-visitor cap bounds one abuser, a global cap bounds the bill.

06
Assert on behaviour, not on wording, in the smoke tests Reversed

Instead of: Match the expected phrasing of a good answer

Two versions failed correct refusals for using different words. A test that checks vocabulary fails the good runs and passes the bad ones.

The reusable playbook

01
Design the refusals first

What a system will not do is a product decision, not a safety afterthought. Write the refusals before the capabilities and the rest of the design arranges itself around them. On a portfolio, the refusal is the credibility.

02
Measure before you architect

Retrieval is the reflex answer to "chat with my documents" because most corpora do not fit in context. Count yours before assuming it doesn't. The sophisticated architecture is sometimes the worse one, and measuring is the only way you find out.

03
Model output is untrusted input

Text a model produced, rendered into a browser, is an injection surface, no different from text a stranger typed. Parse it, escape it, and write the test that tries to get out. My first link pattern allowed protocol-relative URLs. The test caught it, not me reading the code.

04
Pin every fact you cannot afford to have invented

Grounding narrows where a model fills gaps, it does not stop it. The dangerous fabrication is never the absurd one, it is the plausible one shaped exactly like the truth. Name those facts explicitly and test them on every run.

Every other case study here is about designing something for someone else. This one is the same discipline turned on my own site: a real architectural call, a cost I had to justify, a privacy promise written into code rather than into a paragraph, and a bug a test caught before a visitor did. It is running now, at the bottom right of this page. Ask it something it cannot answer. It is quite good at admitting that.

Ask the archive

Grounded in this site’s case studies and journal. It says when it doesn’t know. Questions are recorded so Loïc can see what the site fails to answer. Nothing else is stored.

About this assistant ↗