Nester

2026

NesterAI

Designing the boundary between what an AI answers live and what a human teaches it

Role

Product Designer & Product Manager

Timeline

2026 — in Design

Team

1 Designer/PM · CTO

Platform

Conversational assistant embedded in Nester, consultive/read-only, with a back-office feedback loop that "trains" the agent without retraining the model

Role

Product Designer & Product Manager

Timeline

2026 — in Design

Team

1 Designer/PM · CTO

Platform

Conversational assistant embedded in Nester, consultive/read-only, with a back-office feedback loop that "trains" the agent without retraining the model

Context

A pivot toward proactive AI, tested on read-only ground first.

Nester pivoted in July 2026 toward a "Proactive AI" strategy. The initial launch runs in dogfooding mode (internal use), then closed beta with selected clients — the goal isn't to ship a complete feature in one shot, but to collect real data on user queries to progressively train and improve the model.

The MVP scope is deliberately narrow: consultive, read-only. The AI doesn't execute actions — it answers natural-language questions about data that already exists in the client's ERP (arrears, contracts, incidents, properties). No task automation yet; first, we need to prove the model actually understands what people are asking.

Challenge

The real problem wasn't the chat UI — it was teaching an agent without retraining it.

Designing every state, not just the happy path

I defined the entry point (an icon in the navbar, gated by a per-organization feature flag), the dynamic header, and above all the waiting states: thinking (2–4s), analyzing (when the answer requires ERP-side calculations, +5s), and long wait/latency (at 12s, a reinforcement message instead of leaving the user staring at an unexplained ellipsis).

Mapping the failure modes before the interface

I mapped every possible edge case before designing the happy path — no results, NLU error, timeout, action not allowed, API down — each with its own message, whether it asks for feedback or not, and what recovery action it offers (a retry button, blocking the input, redirecting to human support).

The question that went to engineering

How do you get an AI agent to "learn" from a human's correction without retraining the model, and without the rules file turning into an unmanageable mess three months in? I brought that question into a joint technical review with the CTO, and the outcome directly changed how I designed the back-office.

Approach

Two systems, one question: what does the AI compute live, and what does a human decide?

Conversational states built around real questions

The suggestion chips on the initial state ("Which tenants are in arrears?", "Contracts expiring in 60 days", "Urgent open incidents", "Today's summary") aren't generic examples — they're the queries the product team identified as highest-value for a property manager's actual day-to-day.

Separating business facts from behavior rules

My initial proposal was a rules file where every human correction translated into a text entry. In the technical review with the CTO, we identified a deeper problem: we were mixing two very different types of knowledge under the same document — business facts (which reports a client can see, changing per client in real time) and behavior rules (tone, formality — relatively stable). Mixing them is the direct cause of a rules file growing out of control. The fix: business facts get resolved through tool calling, not a static rule; behavior rules stay in a versioned file, categorized by theme (tone/style, product/business, chat behavior, restrictions) so conflict detection stays tractable.

Open decisions brought to the table, not assumed

I didn't unilaterally decide whether feedback gets captured per individual response or per full conversation, or whether chat features vary by contracted plan (tiers) from day one. I left them explicit as pending product/business decisions, with a recommended simple starting point and a clear path to scale if needed.

Talking point: this is the part I want to emphasize most in the interview — I didn't design screens and throw them over the wall to engineering. I was part of the architecture conversation, and that directly changed how I designed the review back-office.

Solution

A chat interface and a review back-office, designed as one system.

The assistant's interface

Three main states: Landing/Home (suggestion chips + input), Chat (with differentiated loading sub-states), and History (searchable, grouped chronologically, with individual deletion). The main input and the footer input share the same submission logic.

The edge-case table

Five failure scenarios defined one by one (no results, NLU error, timeout, action not allowed, API down), each with its exact copy, whether it requests user feedback, and its specific recovery action.

The review back-office

A master table of conversations (filterable by Pending/Reviewed, searchable, with tokens and latency visible per row) and a detail modal where the internal team flags a response as incorrect, states the error reason, and leaves a technical comment for engineering. That feedback stays pinned under the response and feeds directly into the rule-generation system: an LLM proposes a generalized rule from the comment, a human confirms it before it's inserted, and only then does it move to production with versioning and rollback available.

Reflection & what's next

Still in Design — and that's the point of showing it here.

This project hasn't been built yet — which is exactly why it's the case study that lives only in the interview, not in the public portfolio. The conversation with the CTO left a concrete list of open decisions before building (tiers from day one or not, feedback granularity, which evaluation tool to adopt, where to set the golden-set cutoff). Showing that unresolved decision-making process, instead of only a polished final result, is a more honest picture of what real AI product design work looks like. The rollout is phased on purpose: internal dogfooding first, closed beta next, progressive activation by client subset after — deliberately deferred until there's enough conversation volume for the feedback signal to be statistically reliable.

Key Takeaway

Designing an AI product isn't designing the chat interface.

The part of the design work that mattered most here was drawing the line between what the system resolves live and what a human decides — and how that decision becomes durable knowledge. That line doesn't get drawn alone in Figma; it gets drawn in a conversation with engineering about which technical pattern fits which type of problem.

Josefina Yost

Product Designer with a PM's mindset — complex problems, complete solutions

Contact

josefinayost@gmail.com

Josefina Yost

Product Designer with a PM's mindset — complex problems, complete solutions

Contact

josefinayost@gmail.com

Josefina Yost

Product Designer with a PM's mindset — complex problems, complete solutions

Contact

josefinayost@gmail.com