About Ulvi
Ulvi is a character-based AI app that people talk to every day. Every character has a voice, an environment, and a capability of their own: one reads dreams, one pulls tarot, one runs breathing sessions, one helps you sort a messy decision, and new characters bring new capabilities. We launched in August 2026 and we are building for a global audience, with Türkiye and Brazil as our first focus markets. The team is small, so what you ship goes live quickly and you get to watch people actually use it.
About the role
You read what the AI actually said, decide whether it was good, and know exactly what to change so the next one is better. Elsewhere this job is called LLM quality analyst, conversation designer or AI trainer. Here, it is half editor, half analyst.
Team
You join the Product Success team and report directly to the founders. You work most closely with our Character Writer: they write the characters, you measure how they hold up in the wild.
What you will do
- Define what “good” means for a conversation, in criteria specific enough that someone can disagree with them, and sturdy enough to survive a much bigger cast in three languages.
- Read a lot of conversations: some are sessions with a clear shape, some are long chats about nothing. Find where the tone drifted, where the pacing died, and what people keep asking for.
- Catch the failures that sound good: an invented source, a borrowed voice, a confident answer that missed the point.
- Improve prompts and measure whether the change helped or only felt like it did. The work lives in logs, spreadsheets and prompts: no code required, no fear of it either.
- Characters have a canon someone else wrote. You notice when they drift from it, report where, and propose the fix.
- Turn what you find, in the conversations and in user feedback, into short notes the team can act on the same week.
What we need
- Hands-on curiosity about LLMs. You know what a system prompt is, and you have poked at models to see where they break; ideally you have run one on your own machine. The fixes in this job happen inside the prompt, not beside it.
- Enough rigour to tell an improvement from a coincidence. When a prompt change seems to work, you check whether it holds across many conversations or just looked good in three lucky ones.
- A feel for what makes a session land. Each capability gives its sessions a different shape: a dream reading, a tarot pull and a breathing session each flow differently. You can tell when one carries and when it drags, and say why in a sentence someone can act on.
- Language judgement that holds in Turkish and English. You hear the difference between warm and saccharine, direct and cold, funny and trying too hard. Criteria, notes and other working documents are written in English; most of the conversations are not.
- Patience for reading volume. You still notice the problem in the 400th conversation after reading 399 fine ones.
- Tracing and observability. We run our conversations through Langfuse. You have used it or a similar tracing tool, or you understand what tracing does and are eager to get your hands on it. This is technical quality assurance built on editorial judgement, not content editing under another name.
Nice, not required
- Portuguese. We run in three languages, and someone who can judge all three covers far more of the product than someone who covers one.
- Editorial, translation, localization or linguistics background: any experience where you had to judge language quality and defend the judgement.
- Enthusiasm for working across very different topics: tarot one month, breathing the next, something nobody predicted after that.
How we work
- Fully remote, from anywhere in Türkiye.
- Full-time. Working hours are 10:00–18:00 (TR).
- Turkish and English day to day: Turkish in the conversations, English in the documents.
- Written-first: short status notes, decisions in documents, few meetings.
- Bring your own device: we do not provide hardware.
- Communication runs on Discord.
How hiring works
- We care about what you can do more than where you learned it. Nobody here will ask which university you went to, or whether you went at all.
- We read everything that comes in and shortlist from CVs and portfolios.
- We share the salary range, based on your CV and experience, before the first interview.
- If we shortlist you, we reach out with a small task. It is short, it looks like the actual work of the role, and it is different for each role.
- Depending on the task, we will ask you either to send it in ahead or to present it in the interview. If your portfolio already covers the same ground, that can count too.
How to apply
Send your application to hello@ulvi.ai and attach:
- Your CV
- A cover letter
- A portfolio, work samples, or any other material you believe shows what you can do