How can we reduce the time it takes for marketers to experiment with their ideas?
I replaced a 47-field campaign form with a conversation.

How can we reduce the time it takes for marketers to experiment with their ideas?
I replaced a 47-field campaign form with a conversation.

How can we reduce the time it takes for marketers to experiment with their ideas?
Designing the mental model for AI-native campaign creation.

Plotline | 2026

Plotline | 2026

We added AI to every field in the product. Then we deleted the fields and changed the whole workflow.

Conversational AI for consumer apps //
Can AI actually help users in-context?

Plotline helps growth teams at consumer apps build in-app engagement - stories, nudges, widgets, quizzes, scratch cards, spin wheels - without code. By 2024, the product worked well. Marketers could build anything. The question was whether they knew what to build.

This case study covers two phases of work.

In the first, I added AI to the existing campaign builder - text suggestions, smart scheduling, audience builder, an inspiration gallery. We basically added AI native capabilities to every building block of an in-app campaign. It made parts of the workflow maybe 20% faster.

In the second, I scrapped the campaign builder and replaced it with a conversational agent that works alongside the marketer from the first moment of "what should I even do here?" to the campaign going live. The second phase came out of research that showed me the first phase was solving the wrong problem. Most of this case study is about that research and what it led to.

This is the story of that redesign, the bets it required, and the system it became.

Plotline helps growth teams at consumer apps build in-app engagement - stories, nudges, widgets, quizzes, scratch cards, spin wheels - without code. By 2024, the product worked well. Marketers could build anything. The question was whether they knew what to build.

This case study covers two phases of work.

In the first, I added AI to the existing campaign builder - text suggestions, smart scheduling, audience builder, an inspiration gallery. We basically added AI native capabilities to every building block of an in-app campaign. It made parts of the workflow maybe 20% faster.

In the second, I scrapped the campaign builder and replaced it with a conversational agent that works alongside the marketer from the first moment of "what should I even do here?" to the campaign going live. The second phase came out of research that showed me the first phase was solving the wrong problem. Most of this case study is about that research and what it led to.

This is the story of that redesign, the bets it required, and the system it became.

Plotline's current tools allowed teams to create in-app experiences like these

Team

Team

1 Designer, 1 PM, 5 Engg

1 Designer, 1 PM, 5 Engg

From scoping to launch

From scoping to launch

12-14 weeks

12-14 weeks

Time to live campaign

Time to live campaign

~40%

~40%

Reduction in time-to-value

Reduction in time-to-value

~20%

~20%

Reduction in drop-offs

Reduction in drop-offs

~15-20%

~15-20%

Avg reduction in support tickets

Avg reduction in support tickets

~30-40%

~30-40%

After image

PRE-CURSOR // THE FIRST ROUND OF AI ASSISTS WE DESIGNED IN 2024

What do marketers do with Plotline?
and how do we make it easier

I started where everyone was starting in 2024. The campaign creation flow had a lot of fields - audience rules, channel selection, creative configuration, copy, scheduling, goals. I went through each step and asked where AI could cut time.

AI Assist 1: Copy suggestions

Every text input got an "Ideate with Plotline AI" option. Click it, get three options tuned to your audience and goal, pick one or riff on it. I used the word "Ideate" instead of "Generate" on purpose - small thing, but "generate" makes it sound like the machine is doing your job, "ideate" keeps it collaborative. That instinct turned out to matter a lot more in the next phase.

AI Assist 2: Brand traits

This came directly from research. I talked to about twenty growth teams across gaming, fintech, and e-commerce - Jar, Niyo, Khatabook, others. The pattern I kept hearing wasn't "AI is too slow" or "AI isn't good enough." It was "AI will make us sound like everybody else." The growth team at Niyo was particularly blunt about it - they're a young fintech brand and that identity shows up in every message they send. They weren't comfortable handing that over to a model.

"We want to be very conscious of how we are using AI to generate campaigns because we don't want to lose our brand identity and uniqueness "

Growth team, Kredivo (Indonesia's top 5 finance company)

AI Assist 3: Audience & scheduling assist

Timing matters a lot in lifecycle marketing - nudge someone at the wrong moment and you've burned a touch for nothing. I built recommendations based on the audience's behavioral patterns and goal criteria. Not just "send at 10am" but specific windows backed by actual usage data

AI Assist 4: Recipes

An inspiration gallery for marketers who knew their goal and audience but were blank on what to actually build. Campaign templates cross-referenced by industry and use case, most of the configuration pre-filled, one click to launch.

Plotline's current tools allowed teams to create in-app experiences like these

These features shipped, they got used, they did what they were supposed to do. But here's the thing, and this took me a few months to fully register: a marketer using Plotline after V1 still opened the same campaign builder, still navigated the same screens, still filled in essentially the same fields. The workflow was identical. I'd put better tools inside an unchanged process. Thus, this was just an incremental update and didn't fully capture the changing mental models of marketers.

Who is getting affected

Who is getting affected

Marketers - not able to experiment with the winning strategy with confidence
Product owners are not able to convert their users and plug holes in business growth

Why does it matter

Why does it matter

For a user applying for a loan or making a first investment, that wait is the conversion killer.
Dropped-off users tend to raise support tickets leading to overhead on the business team and delayed clarity

REALISATION // WHY THIS WASN'T ENOUGH AND FINDING THE REAL PROBLEM

User interviews

User interviews

User interviews

To understand their challenges better, we started 1:1 semi-structured interviews which gave us insights into how the growth teams functioned across different organisations.

  • Describe a day in your life as a growth/product manager.

  • Can you walk us through your process of launching a campaign? (conceptualising, configuring and analysing)

  • What are the building blocks of a campaign? How do you formulate them?

  • What would be your ideal experience of improving growth metrics for your users?

  • How can Plotline help you further?

  • What would be some of your future use cases - single/multi channel?

To understand their challenges better, we started 1:1 semi-structured interviews which gave us insights into how the growth teams functioned across different organisations.

  • Describe a day in your life as a growth/product manager.

  • Can you walk us through your process of launching a campaign? (conceptualising, configuring and analysing)

  • What are the building blocks of a campaign? How do you formulate them?

  • What would be your ideal experience of improving growth metrics for your users?

  • How can Plotline help you further?

  • What would be some of your future use cases - single/multi channel?

To understand their challenges better, we started 1:1 semi-structured interviews which gave us insights into how the growth teams functioned across different organisations.

  • Describe a day in your life as a growth/product manager.

  • Can you walk us through your process of launching a campaign? (conceptualising, configuring and analysing)

  • What are the building blocks of a campaign? How do you formulate them?

  • What would be your ideal experience of improving growth metrics for your users?

  • How can Plotline help you further?

  • What would be some of your future use cases - single/multi channel?

Primary research

Primary research

Primary research

I conducted interviews with leading consumer apps from across gaming, finance and e-commerce to understand the natural workflow of creating a campaign right from identifying the gap to solve, ascertaining the target audience and finally creating native looking UI elements and launching them.

Here are a few teams I interviewed among 20+ interviews and the key takeaways.

I conducted interviews with leading consumer apps from across gaming, finance and e-commerce to understand the natural workflow of creating a campaign right from identifying the gap to solve, ascertaining the target audience and finally creating native looking UI elements and launching them.

Here are a few teams I interviewed among 20+ interviews and the key takeaways.

I conducted interviews with leading consumer apps from across gaming, finance and e-commerce to understand the natural workflow of creating a campaign right from identifying the gap to solve, ascertaining the target audience and finally creating native looking UI elements and launching them.

Here are a few teams I interviewed among 20+ interviews and the key takeaways.

Three findings I had missed

Three findings I had missed

Three findings I had missed

After V1 shipped, adoption was fine but it wasn't moving the numbers we cared about. I went back to the same transcripts with a different question, ran five additional interviews focused on the pre-creation phase, and pulled session recordings and time-on-task data from V1.

After V1 shipped, adoption was fine but it wasn't moving the numbers we cared about. I went back to the same transcripts with a different question, ran five additional interviews focused on the pre-creation phase, and pulled session recordings and time-on-task data from V1.

The real bottleneck was before the builder.

  • Which nudge would work for whom?

  • How do I ensure that my users have a smooth experience and are not disturbed with nudges?


  • Which nudge would work for whom?

  • How do I ensure that my users have a smooth experience and are not disturbed with nudges?


Marketers think in goals, not campaigns.

  • Every organisation was sure of what cohorts of users to target.

  • This along with the goal metrics were almost frozen because other stakeholders are also involved

  • Every organisation was sure of what cohorts of users to target.

  • This along with the goal metrics were almost frozen because other stakeholders are also involved

Alignment was the hidden time sink.

  • They don't get enough time to iterate and come up with creative solutions to their problems

  • Have little or no time left for research on current market trends and rely on historical approach

  • They don't get enough time to iterate and come up with creative solutions to their problems

  • Have little or no time left for research on current market trends and rely on historical approach

Their ideal experience

  • Brand-safe results that I can experiment with quickly

  • A tool which lets them know what could help them understand what is working and for whom

  • Brand-safe results that I can experiment with quickly

  • A tool which lets them know what could help them understand what is working and for whom

Before V1, I ran interviews with 20+ growth and marketing leads across gaming, fintech, and e-commerce — teams at Jar, Niyo, Khatabook, Winzo, and others. The goal was straightforward: understand where AI could reduce friction in campaign creation.

The interview structure was a mix of contextual inquiry and semi-structured conversation. I'd ask them to walk me through a recent campaign end-to-end — not hypothetically, but pulling up their actual screens, showing me what they did and where they got stuck. Then I'd dig into the sticking points.

The findings from this first round gave me V1:

Marketers spent the most visible time on copy and creative iteration. They'd write a headline, second-guess it, rewrite it, check it against their brand voice, try again. AI text suggestions were the obvious intervention.

Scheduling was largely guesswork. Most teams sent campaigns at round-number times (10am, 2pm) because they didn't have behavioral data surfaced at the point of decision. Scheduling recommendations had a clear opening.

Campaign ideation was scattered. When a marketer didn't know what to build, they'd browse competitors, scroll through old campaigns, ask teammates. Recipes consolidated that into one place.

These were real pain points and the V1 features addressed them. But here's what I missed in this first round of research: I was watching what happened inside the builder. I wasn't paying enough attention to what happened before the builder was even opened.

Before V1, I ran interviews with 20+ growth and marketing leads across gaming, fintech, and e-commerce — teams at Jar, Niyo, Khatabook, Winzo, and others. The goal was straightforward: understand where AI could reduce friction in campaign creation.

The interview structure was a mix of contextual inquiry and semi-structured conversation. I'd ask them to walk me through a recent campaign end-to-end — not hypothetically, but pulling up their actual screens, showing me what they did and where they got stuck. Then I'd dig into the sticking points.

The findings from this first round gave me V1:

Marketers spent the most visible time on copy and creative iteration. They'd write a headline, second-guess it, rewrite it, check it against their brand voice, try again. AI text suggestions were the obvious intervention.

Scheduling was largely guesswork. Most teams sent campaigns at round-number times (10am, 2pm) because they didn't have behavioral data surfaced at the point of decision. Scheduling recommendations had a clear opening.

Campaign ideation was scattered. When a marketer didn't know what to build, they'd browse competitors, scroll through old campaigns, ask teammates. Recipes consolidated that into one place.

These were real pain points and the V1 features addressed them. But here's what I missed in this first round of research: I was watching what happened inside the builder. I wasn't paying enough attention to what happened before the builder was even opened.

Were the AI assists we added a vitamin or a painkiller?

The thing that kept coming up: by the time a marketer opened the campaign builder, they'd already done all the hard thinking. They'd already noticed a metric moving wrong, already debated with their team what might help, already decided whether it was worth testing. The builder was just the last step: the execution. And execution was the part they were already fast at.

The thing that kept coming up: by the time a marketer opened the campaign builder, they'd already done all the hard thinking. They'd already noticed a metric moving wrong, already debated with their team what might help, already decided whether it was worth testing. The builder was just the last step: the execution. And execution was the part they were already fast at.

Our AI features were making the fast part faster. The slow part, figuring out what to build and whether it would work happened entirely outside Plotline, usually in a mix of analytics tools, Slack threads, past campaign data, and gut feel. No assists involved or at least no translation to our tool.

Our AI features were making the fast part faster. The slow part, figuring out what to build and whether it would work happened entirely outside Plotline, usually in a mix of analytics tools, Slack threads, past campaign data, and gut feel. No assists involved or at least no translation to our tool.

I went back to my interview transcripts. Instead of "where do they get stuck in the builder?" I was looking for "what happens before they open the builder?"

I went back to my interview transcripts. Instead of "where do they get stuck in the builder?" I was looking for "what happens before they open the builder?"

The pattern was pretty consistent:

They notice something. A metric drops, a seasonal moment approaches, a cohort starts behaving differently. Then they form a rough hypothesis — "we should probably re-engage these users" or "this looks like it'd respond to an education-led approach." Then they evaluate whether it's worth the effort — can I get buy-in? what's the opportunity cost? Then, finally, they build.

Our product entered the picture at step four. Steps one through three were on their own.

The pattern was pretty consistent:

They notice something. A metric drops, a seasonal moment approaches, a cohort starts behaving differently. Then they form a rough hypothesis — "we should probably re-engage these users" or "this looks like it'd respond to an education-led approach." Then they evaluate whether it's worth the effort — can I get buy-in? what's the opportunity cost? Then, finally, they build.

Our product entered the picture at step four. Steps one through three were on their own.

  1. They notice something.

A metric drops, a seasonal moment approaches, a cohort starts behaving differently.

  1. Then they form a rough hypothesis

"We should probably re-engage these users" or "This problem looks like it'd respond to an education-led approach."

  1. Then they evaluate whether it's worth the effort

Can I get buy-in? what's the opportunity cost? Then, finally, they build.

Plotline's current tools allowed teams to create in-app experiences like these

“Over the past 2 years, we have seen that the time taken to resolve support tickets is inversely proportional to lifetime value of our customers"

ALETHIA TAN

ALETHIA TAN

SVP, Growth, Kredivo Indonesia

SVP, Growth, Kredivo Indonesia

The pattern was consistent across every fintech customer we talked to. The gap wasn't information - the information existed somewhere, in a policy doc, or in an FAQ buried three taps deep. The gap was timing and context. The right information, but not available at the moment the user needed it, in the place they were looking for it

STARTING OVER // ONCE I SAW IT THIS WAY, I COULDN'T UNSEE IT

We talked to the people losing these users

What we didn't know

What we didn't know

Before designing anything, I needed to understand what "answering user questions in real-time" actually meant to the people who'd be responsible for it - product and growth teams at our customer companies.

  1. Where do drop-offs actually happen, and which of those moments could a conversation genuinely help?

  2. What would make a marketer trust an AI agent enough to deploy it to their users?

  3. What do end users expect from an AI inside a financial app - and where does trust break down?

  4. What does "good" look like for a conversational interaction in a high-stakes context (lending, investments)?

How we found the answers

How we found the answers

We ran 1:1 semi-structured interviews with product and growth leads at fintech apps in our customer base. Semi-structured because we wanted to follow threads, we had a guide, but the most valuable findings came from places we didn't expect. We talked to teams at Kredivo, Dream11, and several others across lending, investing, and wallets.

We also did a journey audit, mapped the core user flows (loan application, investment, wallet top-up/withdrawal) against where support tickets were being raised. This gave us a quantitative layer to ground the qualitative interviews.

One deliberate gap: no end-user research upfront. The timeline didn't allow it. We'd validate with real users during the pilot. That was a tradeoff we named out loud

  • What are the main customer interactions within your app that could benefit from AI-powered conversations (e.g., customer support, product recommendations, order tracking)?

  • How do you currently gather customer feedback, troubleshoot issues, and upsell products? Would a conversational agent be suitable for any of these?

  • How do you estimate the agent's impact in your app? (e.g., multilingual support, personalization, deep product knowledge)?

  • How important is AI-human handover in complex cases? What is your expectation of bot vs. human interactions?

  • What are your top concerns about integrating conversational AI agents? (Options: technical complexity, security and privacy, customer trust, handling edge cases, impact on brand, regulatory compliance)

“Over the past 2 years, we have seen that the time taken to resolve support tickets is inversely proportional to lifetime value of our customers"

ALETHIA TAN

ALETHIA TAN

SVP, Growth, Kredivo Indonesia

SVP, Growth, Kredivo Indonesia

“An agentic experience inside my app should aid the overall discoverabilty and usage. It should intelligently understand when it is needed and what it should help with”

Rishabh

Rishabh

Growth team, Dream11

Growth team, Dream11

Three things we learned that reshaped our direction

Three things we learned that reshaped our direction

We expected the dominant concern to be "Will the AI give wrong answers?" That was a concern, but it wasn't the primary one.

Training: "How do I make it know what it needs to know?"

Training: "How do I make it know what it needs to know?"

Marketers were anxious about knowledge gaps - stale information, missing context, wrong answers. But what surprised us was how they wanted to solve this. They wanted to teach the system from conversations they'd already had.

This insight directly shaped the evaluation through a benchmark system.

Testing: "How do I know it'll behave the right way before I push it live?"

Testing: "How do I know it'll behave the right way before I push it live?"

There was near-universal anxiety about deploying something they couldn't fully preview. Teams wanted to simulate conversations - not just check settings. This wasn't about technical QA. It was about confidence. A marketer needs to be able to say "I have talked to this thing and it makes sense" before they trust it with their users.

Deployment: "What happens when it doesn't know something or gets it wrong?"

Deployment: "What happens when it doesn't know something or gets it wrong?"

The question of escalation - when does the bot hand off to a human, and how - came up in every single interview. More importantly, several people raised brand risk: "If my AI agent says something incorrect about a loan product, I'm liable." This wasn't paranoia. It was valid. It changed how we thought about the autonomy spectrum.

APPROACH // WHAT A SOLUTION WOULD NEED TO DO

Key requirements that narrowed the field

Key requirements that narrowed the field

Coming out of research, the shape of the solution was getting clearer. Whatever we built needed to solve for these.

  1. Know where the user is in the app, not just "on the home screen" but "midway through a loan application, on the documentation step."
  1. Know what they're trying to do and tailor the response to that specific intent, not a generic FAQ answer.
  1. Handle questions it hasn't been specifically programmed for the organic, contextual, long-tail questions that no scripted system can anticipate.
  1. Respond in real-time because a 1–2 hour support ticket is a risky approach as it may lead to user abandonment.
  1. Know when to stop because when a question touches compliance, liability, or something it genuinely doesn't know, it needs to hand off or escalate.

Why the obvious options didn't work

The risks were real: hallucination, inconsistency, latency, brand voice drift. But these were design problems, not reasons to abandon the approach. Every structural decision in the product - the knowledge base architecture, the benchmark system, the testing simulator, the deployment controls - exists to constrain and direct the LLM, not to replace it.

Decision trees / scripted flows

Decision trees / scripted flows

Decision trees are deterministic and auditable, but they fail requirement three immediately. You can't pre-build branches for "what happens to my lock-in if I want to prepay in 6 months?" The organic tail is infinite.

Decision trees are deterministic and auditable, but they fail requirement three immediately. You can't pre-build branches for "what happens to my lock-in if I want to prepay in 6 months?" The organic tail is infinite.

FAQ overlays / static knowledge surfaces

FAQ overlays / static knowledge surfaces

Plotline's nudge toolkit already did a version of this. Research told us users had outgrown it. Fails requirement three again, and partially fails requirement one (no awareness of where the user is in their flow).

Plotline's nudge toolkit already did a version of this. Research told us users had outgrown it. Fails requirement three again, and partially fails requirement one (no awareness of where the user is in their flow).

LLM-based responses

LLM-based responses

Flexible, can synthesise context from multiple sources, can hold a conversation across turns, can adapt to tone and intent. An LLM at its core, but wrapped in enough structure that a marketer could configure it, test it, trust it, and ship it.

Flexible, can synthesise context from multiple sources, can hold a conversation across turns, can adapt to tone and intent. An LLM at its core, but wrapped in enough structure that a marketer could configure it, test it, trust it, and ship it.

The risks were real: hallucination, inconsistency, latency, brand voice drift. But these were design problems, not reasons to abandon the approach.

The risks were real: hallucination, inconsistency, latency, brand voice drift. But these were design problems, not reasons to abandon the approach. Every structural decision in the product - the knowledge base architecture, the benchmark system, the testing simulator, the deployment controls - exists to constrain and direct the LLM, not to replace it.

Context

Context

Who the agent is, how it speaks, what it can and can't discuss.

Who the agent is, how it speaks, what it can and can't discuss.

Knowledge

Knowledge

What it knows, from which sources, scoped to which flows.

What it knows, from which sources, scoped to which flows.

Tools & Actions

Tools & Actions

What it can do beyond talking and how would it get information

What it can do beyond talking and how would it get information

Deployment

Deployment

When it appears, who sees it, and how it fails gracefully.

When it appears, who sees it, and how it fails gracefully.

TRAINING YOUR AGENT

How do you teach an AI what it should know (and only what it should know)?

How do you teach an AI what it should know (and only what it should know)?

  1. Knowledge base

Collection of data points that the agent can retrieve as required such as FAQs, policy PDFs, product specs. The raw material for your agent to get started.

Collection of data points that the agent can retrieve as required such as FAQs, policy PDFs, product specs. The raw material for your agent to get started.

  1. Benchmark conversations

Curated Q&A pairs that set the standard for how the agent should respond. Not just "here's the information" but "here's what a good answer sounds like."

Curated Q&A pairs that set the standard for how the agent should respond. Not just "here's the information" but "here's what a good answer sounds like."

  1. Knowledge base

Collection of data points that the agent can retrieve as required such as FAQs, policy PDFs, product specs. The raw material for your agent to get started.

Collection of data points that the agent can retrieve as required such as FAQs, policy PDFs, product specs. The raw material for your agent to get started.

Problem to design for

The risks were real: hallucination, inconsistency, latency, brand voice drift. But these were design problems, not reasons to abandon the approach. Every structural decision in the product - the knowledge base architecture, the benchmark system, the testing simulator, the deployment controls - exists to constrain and direct the LLM, not to replace it.

A knowledge base that ingests everything and scopes nothing is dangerous. It hallucinates confidently from irrelevant sources. The design problem isn't "how do you give the agent more information?" It's "how do you make it reach for the right information at the right moment and ignore the rest?"

The risks were real: hallucination, inconsistency, latency, brand voice drift. But these were design problems, not reasons to abandon the approach. Every structural decision in the product - the knowledge base architecture, the benchmark system, the testing simulator, the deployment controls - exists to constrain and direct the LLM, not to replace it.

Solution

The risks were real: hallucination, inconsistency, latency, brand voice drift. But these were design problems, not reasons to abandon the approach. Every structural decision in the product - the knowledge base architecture, the benchmark system, the testing simulator, the deployment controls - exists to constrain and direct the LLM, not to replace it.

  • Easy addition of documents/URLs with guided flows

  • A marketer can tag a source to a specific user flow. "This document should only be referenced when the user is in the loan application." Scoping reduces hallucination by limiting what the model can reach for in any given moment.

The risks were real: hallucination, inconsistency, latency, brand voice drift. But these were design problems, not reasons to abandon the approach. Every structural decision in the product - the knowledge base architecture, the benchmark system, the testing simulator, the deployment controls - exists to constrain and direct the LLM, not to replace it.

Specific addition of URL based content for only relevant information

Adding usage context with your knowledge documents

  1. Benchmark conversations

Curated Q&A pairs that set the standard for how the agent should respond. Not just "here's the information" but "here's what a good answer sounds like."

Curated Q&A pairs that set the standard for how the agent should respond. Not just "here's the information" but "here's what a good answer sounds like."

Problem to design for

The risks were real: hallucination, inconsistency, latency, brand voice drift. But these were design problems, not reasons to abandon the approach. Every structural decision in the product - the knowledge base architecture, the benchmark system, the testing simulator, the deployment controls - exists to constrain and direct the LLM, not to replace it.

Your brand voice is built through the way you interact with your users. Marketers and growth managers were very conscious of preserving their brand voice


This came directly from the research finding about invisible failure - marketers wanted to train the system from conversations that had already happened, not only by uploading PDFs and hoping for the best. They wanted the system to sound like the best of their manual responses and preserve the overall quality of responses.

The risks were real: hallucination, inconsistency, latency, brand voice drift. But these were design problems, not reasons to abandon the approach. Every structural decision in the product - the knowledge base architecture, the benchmark system, the testing simulator, the deployment controls - exists to constrain and direct the LLM, not to replace it.

Solution

The risks were real: hallucination, inconsistency, latency, brand voice drift. But these were design problems, not reasons to abandon the approach. Every structural decision in the product - the knowledge base architecture, the benchmark system, the testing simulator, the deployment controls - exists to constrain and direct the LLM, not to replace it.

  • The benchmark system works like this: a marketer adds a conversation pair - a real theme or topic and the ideal response. The agent calibrates against these.

  • In the simulator, the marketer rates the agent's actual response against the benchmark. Low-rated responses become training signal.

  • Every design decision in this interface (the rating mechanism, the side-by-side comparison, the feedback input) was tested against one question: would a growth lead at Kredivo/SBI/BharatPe/Upstox understand what they're doing here?

The risks were real: hallucination, inconsistency, latency, brand voice drift. But these were design problems, not reasons to abandon the approach. Every structural decision in the product - the knowledge base architecture, the benchmark system, the testing simulator, the deployment controls - exists to constrain and direct the LLM, not to replace it.

The benchmark interface frames this as teaching by example. Not weights. Not fine-tuning. Not parameters. Just: here's what good looks like. The agent moves toward those responses over time.
The benchmark interface frames this as teaching by example. Not weights. Not fine-tuning. Not parameters. Just: here's what good looks like. The agent moves toward those responses over time.

Benchmark addition for human-like responses in your authentic brand voice

Blind testing via side-by-side rating mechanism compares agent responses to previous benchmarks. This helps give unbiased directions to the system.
Blind testing via side-by-side rating mechanism compares agent responses to previous benchmarks. This helps give unbiased directions to the system.

Blind rating system for actual agent responses

CONFIGURING YOUR AGENT

How do you make "configure an AI" feel like a job a marketer already knows how to do?

How do you make "configure an AI" feel like a job a marketer already knows how to do?

Setting up your agents

Setting up your agents

Since the whole concept of having an AI agent take care of your users' needs, aspirations and frustrations was new, I built a few pre-configured templates to help the marketers get started and explore in a low friction way.

Since the whole concept of having an AI agent take care of your users' needs, aspirations and frustrations was new, I built a few pre-configured templates to help the marketers get started and explore in a low friction way.

Since the whole concept of having an AI agent take care of your users' needs, aspirations and frustrations was new, I built a few pre-configured templates to help the marketers get started and explore in a low friction way.

To reduce the cognitive load on the marketer, I designed a template library, pre-configured setups for common use cases. A loan FAQ agent. An investment onboarding assistant. A wallet support agent. They start from something recognisable and shape it.
To reduce the cognitive load on the marketer, I designed a template library, pre-configured setups for common use cases. A loan FAQ agent. An investment onboarding assistant. A wallet support agent. They start from something recognisable and shape it.

Pre-configured agent templates as the starting block. These can be contextual to any industry the dashboard is configured for

Once a marketer picks a template, they're configuring the agent's behaviour. The risk here was dumping everything into one long settings panel - tone, rules, boundaries, flow logic - and hoping they'd figure it out. Thus, I broke context into three blocks, each with a distinct job:

Communication style, Conversation flow and Escalation rules

The three blocks create a natural sequence: first you decide how the agent sounds, then how it navigates, then where it draws the line. Each block is completable independently, and each has sensible defaults that work out of the box.

Once a marketer picks a template, they're configuring the agent's behaviour. The risk here was dumping everything into one long settings panel - tone, rules, boundaries, flow logic - and hoping they'd figure it out. Thus, I broke context into three blocks, each with a distinct job:

Communication style, Conversation flow and Escalation rules

The three blocks create a natural sequence: first you decide how the agent sounds, then how it navigates, then where it draws the line. Each block is completable independently, and each has sensible defaults that work out of the box.

Broken down context for a clear set of instructions to the system

Measuring success for this flow

Measuring success for this flow

It is very important to track the usability of a completely new product added to our core dashboard. Thus, we are closely tracking the performance.

It is very important to track the usability of a completely new product added to our core dashboard. Thus, we are closely tracking the performance.

It is very important to track the usability of a completely new product added to our core dashboard. Thus, we are closely tracking the performance.

Task success rate - Creation

Task success rate - Creation

Currently ~ 57%

Usability support ticket ratio

Usability support ticket ratio

Currently ~ 45%

Time to first value

Time to first value

Currently ~ 8 minutes

TESTING AND BUILDING TRUST IN YOUR AGENT

How do you make a non-technical person confident enough to deploy AI to their users?

How do you make a non-technical person confident enough to deploy AI to their users?

Configuring an agent and trusting it are two different things. A marketer can fill in every setting correctly and still not feel confident enough to ship it to millions of users. The gap isn't knowledge but evidence. They need to see the agent perform before they believe it will.

Configuring an agent and trusting it are two different things. A marketer can fill in every setting correctly and still not feel confident enough to ship it to millions of users. The gap isn't knowledge but evidence. They need to see the agent perform before they believe it will.

Research was unambiguous on this: every team we spoke to wanted to have an actual conversation with the agent before going live. Not review a settings summary. Not check a preview screenshot. Talk to it. Break it. See how it recovers.

Research was unambiguous on this: every team we spoke to wanted to have an actual conversation with the agent before going live. Not review a settings summary. Not check a preview screenshot. Talk to it. Break it. See how it recovers.

Here’s a snapshot of how we can simulate the entire conversation experience, rate previous conversations and help the agent learn exactly how it is supposed to communicate with your users.

Here’s a snapshot of how we can simulate the entire conversation experience, rate previous conversations and help the agent learn exactly how it is supposed to communicate with your users.

Here’s a snapshot of how we can simulate the entire conversation experience, rate previous conversations and help the agent learn exactly how it is supposed to communicate with your users.

"How will I simulate my user's conversation"

"How will I simulate my user's conversation"

"Can I test agent's responses at scale"

"Can I test agent's responses at scale"

"What is causing latency in replies"

"What is causing latency in replies"

ENSURING THE AGENT SHOWS UP AND BEHAVES THE WAY WE WANT IT TO

How will I deploy the system with confidence?
Giving enough context and situation handling directions to the agent

How will I deploy the system with confidence?
Giving enough context and situation handling directions to the agent

Context broken down into communication styles, conversation guidance & escalation and hand-overs

Context broken down into communication styles, conversation guidance & escalation and hand-overs

How will the system detect when to intervene?

How will the system detect when to intervene?

Adding the right tools and knowledge bases - can we reduce the cognitive load here?

Adding the right tools and knowledge bases - can we reduce the cognitive load here?

Setting up for success? How will I define it?

Setting up for success? How will I define it?

MAKING SENSE OF IT ALL

Analytics that feed back into the system, and keeps constantly improving

Analytics that feed back into the system, and keeps constantly improving

Agent performance broken down into actionable intelligence

Agent performance broken down into actionable intelligence

Started by focusing on core metrics such as conversation volume, goal completion rate, and human handoff rate (when users are escalated to live agents). Real-time conversation logs, knowledge and tools performance also help in targeting the agent better.

Started by focusing on core metrics such as conversation volume, goal completion rate, and human handoff rate (when users are escalated to live agents). Real-time conversation logs, knowledge and tools performance also help in targeting the agent better.

Started by focusing on core metrics such as conversation volume, goal completion rate, and human handoff rate (when users are escalated to live agents). Real-time conversation logs, knowledge and tools performance also help in targeting the agent better.

Four metrics. Each one connects to a specific action the marketer can take.

Four metrics. Each one connects to a specific action the marketer can take.

Conversation quality

Conversation quality

Gaps in responses today become training tasks. The marketer adds a benchmark, updates a source, re-tests. The system improves through use, without engineering involvement.

Gaps in responses today become training tasks. The marketer adds a benchmark, updates a source, re-tests. The system improves through use, without engineering involvement.

Goal completion rate

Goal completion rate

Did the user continue their journey after the conversation ended? This is the metric that ties agent performance to the business outcome customers actually care about.

Did the user continue their journey after the conversation ended? This is the metric that ties agent performance to the business outcome customers actually care about.

Agent deployment

Agent deployment

Is the agent being triggered at the right instances? If not, deployment rules need adjustment. Is the latency affected in multiple parallel conversations?

Is the agent being triggered at the right instances? If not, deployment rules need adjustment. Is the latency affected in multiple parallel conversations?

Human handoff rate

Human handoff rate

How often is the agent reaching its limits? High rates signal knowledge gaps or over-tight escalation thresholds.

How often is the agent reaching its limits? High rates signal knowledge gaps or over-tight escalation thresholds.

"I want the system to learn from its mistakes and not repeat them"

"I want the system to learn from its mistakes and not repeat them"

Ensuring a system of record for diving deeper and debugging

Ensuring a system of record for diving deeper and debugging

WHAT'S NEXT

What I learnt & where are we taking this next

What I learnt & where are we taking this next

Conceptualising how to build modular agentic experiences for platforms like Plotline, for marketers from leading consumer apps and visualising experience for their end users was a great opportunity to understand and design for:

Building conversational agents in a modular way
Building conversational agents in a modular way

Breaking the whole process into functions such as context, knowledge, tools & actions ensured a very gradual learning curve and progressive complexity.

Segregating global and agent-specific building blocks
Segregating global and agent-specific building blocks

Centralising appearance, communication and brand guidelines reduces the potential for inconsistent experiences

Building trust and traceability into every AI decision
Building trust and traceability into every AI decision

Building unbiased testing and learning flows for agent's training solves for trust at a scale of millions

Where are we taking this next

Where are we taking this next

Working on agentic experiences opens a whole world of possibilities. For the next versions, ideations and concepts have already started!

Improving handover flow to manual agents
Improving handover flow to manual agents

Real-time view into handoffs as they happen for live escalation visibility

In-depth end-user research
In-depth end-user research

Sessions with real users of pilot customers to close the gap we deliberately left open.

Introducing new access points and interactions
Introducing new access points and interactions

Access points like floating buttons, pinned banners, gestures like long hold, bottom swipe can be introduced

The interesting problem in AI product design isn't the model. It's the person sitting in the configuration screen, deciding whether this system knows enough, behaves well enough, and fails reliably enough to represent their brand to their users. That's the interface we built.
The interesting problem in AI product design isn't the model. It's the person sitting in the configuration screen, deciding whether this system knows enough, behaves well enough, and fails reliably enough to represent their brand to their users. That's the interface we built.

A true partner for marketers //
Multi-level AI assistant for marketing engagement

A true partner for marketers //
Multi-level AI assistant for marketing engagement

A true partner for marketers //
Multi-level AI assistant for marketing engagement

Plotline | 2025

Plotline | 2025

In 2024, every B2B SaaS company was adding AI to their product the same way: find a text field, put a sparkle icon next to it, generate some content, and ship a blog post about it. I did that too — and then I watched nobody use it. What followed was a nine-month redesign of the entire campaign creation experience at Plotline, from form-filling to conversation, from campaigns to missions, from a tool that waits for instructions to an agent that shows up with a plan.


This is the story of that redesign, the bets it required, and the system it became.

In 2024, every B2B SaaS company was adding AI to their product the same way: find a text field, put a sparkle icon next to it, generate some content, and ship a blog post about it. I did that too — and then I watched nobody use it. What followed was a nine-month redesign of the entire campaign creation experience at Plotline, from form-filling to conversation, from campaigns to missions, from a tool that waits for instructions to an agent that shows up with a plan.


This is the story of that redesign, the bets it required, and the system it became.

Design ethos // Using AI in your worflows should feel like talking to a partner that helps rather than just witnessing magic happen

Design ethos // Using AI in your worflows should feel like talking to a partner that helps rather than just witnessing magic happen

CONTEXT

CONTEXT

Add both images in Media to enable the comparison.

What does Plotline do?

What does Plotline do?

What does Plotline do?

Plotline helps growth teams at consumer apps with onboarding, activation, adoption and retention use cases by letting them build elements like stories, in-line widgets, floating buttons, spotlights, quizzes, scratch cards and much more - without code. We have built a platform for marketers to enable them to nudge the right user at the right time to perform the intended action

PROBLEM

Our users knew what and whom to solve for, but were unsure of what'll work and why

Our users knew what and whom to solve for, but were unsure of what'll work and why

Our users knew what and whom to solve for, but were unsure of what'll work and why

My objective is to enable marketers across industries and scale to engage their users the best way they can. A major part of this is to make comprehensive campaigns that help the operators with tasks such as segmentation, nudge design and making sure every campaign has that unique flavour for which their brand is known for.

Marketers wanted to know in general what was working for which use cases in their industry

They wanted to get inspired and then create and launch experiments at lightning speed

They felt they need to convince all the stakeholders multiple number of times and building inspired campaigns could help

INITIAL DIRECTION

INITIAL DIRECTION

Do marketers need an assistant?

Do marketers need an assistant?

Do marketers need an assistant?

There are many roles that the marketers have to play to decide - what to send, who to send to, when to send and what to optimize. Thus we felt that building a system that can assist them wherever they needed - be it building a complete campaign or just generating text options - will be something that'll add actual value to their daily processes.

We wanted to understand at a deeper level how growth and marketing teams at consumer business ideate and launch experiments and how we can support them.

We wanted to understand at a deeper level how growth and marketing teams at consumer business ideate and launch experiments and how we can support them.

We wanted to understand at a deeper level how growth and marketing teams at consumer business ideate and launch experiments and how we can support them.

Understanding existing workflows, benchmarking, uncovering any latent aspirations

Primary research

Primary research

Primary research

I conducted interviews with leading consumer apps from across gaming, finance and e-commerce to understand the natural workflow of creating a campaign right from identifying the gap to solve, ascertaining the target audience and finally creating native looking UI elements and launching them.

Here are a few teams I interviewed among 20+ interviews and the key takeaways.

I conducted interviews with leading consumer apps from across gaming, finance and e-commerce to understand the natural workflow of creating a campaign right from identifying the gap to solve, ascertaining the target audience and finally creating native looking UI elements and launching them.

Here are a few teams I interviewed among 20+ interviews and the key takeaways.

I conducted interviews with leading consumer apps from across gaming, finance and e-commerce to understand the natural workflow of creating a campaign right from identifying the gap to solve, ascertaining the target audience and finally creating native looking UI elements and launching them.

Here are a few teams I interviewed among 20+ interviews and the key takeaways.

User interviews

User interviews

User interviews

To understand their challenges better, we started 1:1 semi-structured interviews which gave us insights into how the growth teams functioned across different organisations.

  • Describe a day in your life as a growth/product manager.

  • Can you walk us through your process of launching a campaign? (conceptualising, configuring and analysing)

  • What are the building blocks of a campaign? How do you formulate them?

  • What would be your ideal experience of improving growth metrics for your users?

  • How can Plotline help you further?

  • What would be some of your future use cases - single/multi channel?

To understand their challenges better, we started 1:1 semi-structured interviews which gave us insights into how the growth teams functioned across different organisations.

  • Describe a day in your life as a growth/product manager.

  • Can you walk us through your process of launching a campaign? (conceptualising, configuring and analysing)

  • What are the building blocks of a campaign? How do you formulate them?

  • What would be your ideal experience of improving growth metrics for your users?

  • How can Plotline help you further?

  • What would be some of your future use cases - single/multi channel?

To understand their challenges better, we started 1:1 semi-structured interviews which gave us insights into how the growth teams functioned across different organisations.

  • Describe a day in your life as a growth/product manager.

  • Can you walk us through your process of launching a campaign? (conceptualising, configuring and analysing)

  • What are the building blocks of a campaign? How do you formulate them?

  • What would be your ideal experience of improving growth metrics for your users?

  • How can Plotline help you further?

  • What would be some of your future use cases - single/multi channel?

Key findings from interviews

Key findings from interviews

Key findings from interviews

Marketers think about

  • Which nudge would work for whom?

  • How do I ensure that my users have a smooth experience and are not disturbed with nudges?


  • Which nudge would work for whom?

  • How do I ensure that my users have a smooth experience and are not disturbed with nudges?


What are they confident about

  • Every organisation was sure of what cohorts of users to target.

  • This along with the goal metrics were almost frozen because other stakeholders are also involved

  • Every organisation was sure of what cohorts of users to target.

  • This along with the goal metrics were almost frozen because other stakeholders are also involved

Their biggest frustrations

  • They don't get enough time to iterate and come up with creative solutions to their problems

  • Have little or no time left for research on current market trends and rely on historical approach

  • They don't get enough time to iterate and come up with creative solutions to their problems

  • Have little or no time left for research on current market trends and rely on historical approach

Their ideal experience

  • Brand-safe results that I can experiment with quickly

  • A tool which lets them know what could help them understand what is working and for whom

  • Brand-safe results that I can experiment with quickly

  • A tool which lets them know what could help them understand what is working and for whom

What problems to pick

What problems to pick

What problems to pick

When taking on a workflow improvement for the core product, it helps to organise all the findings into problems and challenges which leads to insights and constraints. New ideas and opportunities are distilled from these groups as a direct result of the research process.

When taking on a workflow improvement for the core product, it helps to organise all the findings into problems and challenges which leads to insights and constraints. New ideas and opportunities are distilled from these groups as a direct result of the research process.

When taking on a workflow improvement for the core product, it helps to organise all the findings into problems and challenges which leads to insights and constraints. New ideas and opportunities are distilled from these groups as a direct result of the research process.

How might we

How might we

How might we

Build trust in our recommendations

Giving the controls to marketers to align the recommendation model to their expectations and brand

Giving the controls to marketers to align the recommendation model to their expectations and brand

Give quick and localised ideas wherever needed

Research showed that marketers seek inspiration or recommendation at different levels - sometimes it could just be ideating on the copy used one of their campaigns or just getting to know the right time to send

Research showed that marketers seek inspiration or recommendation at different levels - sometimes it could just be ideating on the copy used one of their campaigns or just getting to know the right time to send

Give comprehensive and global recommendations too

Marketers sometimes also wanted to get recommendations based on just a use case. In these explorative scenarios we have to assist them in ideating new ideas

Marketers sometimes also wanted to get recommendations based on just a use case. In these explorative scenarios we have to assist them in ideating new ideas

Building trust in our recommendations

Building trust in our recommendations

Building trust in our recommendations

“We don’t want to generate content through Gen AI and sound like everybody else. Niyo is a young fintech brand and we communicate that in every message”

“We don’t want to generate content through Gen AI and sound like everybody else. Niyo is a young fintech brand and we communicate that in every message”

“We don’t want to generate content through Gen AI and sound like everybody else. Niyo is a young fintech brand and we communicate that in every message”

Aditya Bhatt

Aditya Bhatt

Growth team, Niyo

Growth team, Niyo

Ensuring every campaign recommendation is tailored to your brand

To build trust, it was essential to let the marketers feel that their brand voice is being captured in every campaign suggestion Plotline AI gives. Along with this, a concept of general brand traits was introduced to give the text recommendations their brand's unique flavour . This ensures that every campaign is tailor-made to your brand.

To build trust, it was essential to let the marketers feel that their brand voice is being captured in every campaign suggestion Plotline AI gives. Along with this, a concept of general brand traits was introduced to give the text recommendations their brand's unique flavour . This ensures that every campaign is tailor-made to your brand.

To build trust, it was essential to let the marketers feel that their brand voice is being captured in every campaign suggestion Plotline AI gives. Along with this, a concept of general brand traits was introduced to give the text recommendations their brand's unique flavour . This ensures that every campaign is tailor-made to your brand.

By developing a concept of 'Brand traits' all the recommendations are highly tailored to your unique brand voice, thus helping build trust with marketers.

Give quick and local ideas wherever needed

Give quick and local ideas wherever needed

Give quick and local ideas wherever needed

“Most of the times I feel it’s trial & error in ensuring we reach the right audience, with the right content at the right time. I want to get suggestions that will give results for my users”

“Most of the times I feel it’s trial & error in ensuring we reach the right audience, with the right content at the right time. I want to get suggestions that will give results for my users”

“Most of the times I feel it’s trial & error in ensuring we reach the right audience, with the right content at the right time. I want to get suggestions that will give results for my users”

Utkarsh Garg

Utkarsh Garg

Head of Growth, Jar

Head of Growth, Jar

Text specific recommendations

Every text input block while creating campaigns gives you the option to 'Ideate with PlotlineAI'. Subtle UX decisions such as using Ideate in place of Get suggestions or Ask Plotline AI are intended to make the marketer feel that is a collaborative space, and not where answers appear out of magic.

Every text input block while creating campaigns gives you the option to 'Ideate with PlotlineAI'. Subtle UX decisions such as using Ideate in place of Get suggestions or Ask Plotline AI are intended to make the marketer feel that is a collaborative space, and not where answers appear out of magic.

Every text input block while creating campaigns gives you the option to 'Ideate with PlotlineAI'. Subtle UX decisions such as using Ideate in place of Get suggestions or Ask Plotline AI are intended to make the marketer feel that is a collaborative space, and not where answers appear out of magic.

Scheduling recommendations

Scheduling or setting when these campaigns will be delivered to the end users can have a lot of difference in impact. If we nudge the user in the right context, then only chances of conversion are maximised. Thus, based on the audience and goal criteria, we built a system to recommend the 'right time to send'.

Scheduling or setting when these campaigns will be delivered to the end users can have a lot of difference in impact. If we nudge the user in the right context, then only chances of conversion are maximised. Thus, based on the audience and goal criteria, we built a system to recommend the 'right time to send'.

Scheduling or setting when these campaigns will be delivered to the end users can have a lot of difference in impact. If we nudge the user in the right context, then only chances of conversion are maximised. Thus, based on the audience and goal criteria, we built a system to recommend the 'right time to send'.

Generic inspiration - 'Recipes'

To solve for the cases when marketers know what goals they want to achieve by targeting a specific cohort, we came up with the concept of recipes. We cross-referenced the industries and use-cases for a collection of inspiration campaigns.

To solve for the cases when marketers know what goals they want to achieve by targeting a specific cohort, we came up with the concept of recipes. We cross-referenced the industries and use-cases for a collection of inspiration campaigns.

To solve for the cases when marketers know what goals they want to achieve by targeting a specific cohort, we came up with the concept of recipes. We cross-referenced the industries and use-cases for a collection of inspiration campaigns.

With 'Recipes (inspiration gallery)', marketers could launch 1 click campaigns with a majority of UI element selection, content and styling already taken care of!

Comprehensive global recommendations for building campaigns from scratch at hyper-speed

Comprehensive global recommendations for building campaigns from scratch at hyper-speed

Comprehensive global recommendations for building campaigns from scratch at hyper-speed

“Sometimes I feel that quick ideations can not only help us align internally faster on what to launch but also enable us to experiment more. I get a lot of generic references online but it takes time to contextualise and decide whether they would work for my users ”

“Sometimes I feel that quick ideations can not only help us align internally faster on what to launch but also enable us to experiment more. I get a lot of generic references online but it takes time to contextualise and decide whether they would work for my users ”

“Sometimes I feel that quick ideations can not only help us align internally faster on what to launch but also enable us to experiment more. I get a lot of generic references online but it takes time to contextualise and decide whether they would work for my users ”

SANSKRUTI SHARMA

SANSKRUTI SHARMA

Monetisation , Khatabook

Monetisation , Khatabook

How it works

Here’s a snapshot of how the system works. Plotline AI sequentially builds full campaign recommendation for your particular use case and cohort of users as follows

Here’s a snapshot of how the system works. Plotline AI sequentially builds full campaign recommendation for your particular use case and cohort of users as follows

Here’s a snapshot of how the system works. Plotline AI sequentially builds full campaign recommendation for your particular use case and cohort of users as follows

Step 1 - Some quick options like creating a campaign, a feedback survey or an instagram-like story are given to the user

Step 2 - Marketer shares the use case they want to solve

Step 3 - Ideas are generated based on the use case and channels that have worked for this group of users before

Thus, by generating contextual recommendations from scratch helps marketers ideate, iterate and launch experiments faster and in a more precise manner.

Many parts of this project are not revealed here, I'll be happy to go over the complete process and results over a call