0-1 AI-powered

FIntech | Data-heavy |

PitchBook's first AI-powered data collection platform that cut working time and raised data quality and delivery speed

Project Overview

My Role

Lead Product Designer

Product Designer

Product Designer

Team

Cross-functional with PM, AI/ML, Engineering

1x Product Manager

1x Product Manager

Scope

6 month

launched Q1 2026

Achievement

0-1

fresh 成功launch blablh

Product Designer

Product Designer

Deal with ambiguity

不确定scopr

1x Product Manager

1x Product Manager

Simplified complexity

6 month

launched Q1 2026

Context

The most struggle point for investors is to find high-quality,c unique data though big amount of data(海底捞针)and make sure the data is accurate sometimes the 针is wrong. (符合第一个complex人设)

Challenge

But as the datasets grew, hand-collection couldn't keep up. Slow, manual, error-prone. Clients complained about errors and long waits.

GOAL

GOAL 123 business goals+ design goal: reduce internal workload: users able to efficiently update data

PRoblem STATMENT

HMW

Phase 1 research - find the root cause

User interview
1. user quote 图
2. user current flow/joruney map简单的图,画map, 展示daily routine上其他网站搜
Current the biggest struggle/root cause is XXX


Summary of pain points 123
1. user pain point/have to jump across different tools 他们需要去不同平台收集数据)

(最后AI centralized)

Current soltuion audit
从而体现哪些符合user daily需求 哪些是struggling point, 现在的solution 能满足user need, 但是不能满足productivity, (只说现有产品不好) (不用全说) (从problem, research到ideation 都要match)

Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
Dataset Tool
30+ TOOLSfragmented, no shared standard
unified
UDCPUnified Data Collection Platform
CollectionEditor
QC + ReviewSurfaces
Entity Profiles+ Entitlements
AI Extraction + Confidence Scoring Layer
1 PLATFORMshared standards, AI-native, quality enforced

到ideation 开始引入AI,“我们用ai来解决问题” 开始讲collaboration, 因为ai 不是我能做的 (traditiaonlly, 2 decade problem, ai enginner给了很大support) 结合pm stakeholders讨论, 最终的决定/结论是 AI可以帮我们解决这个问题

HOwever, AI is not the perfect soltuion, user no trsut AI, has concern 最重点最重要的分析(核心)
所以我们要定哪几个metrics 从而保证user可以trust, keep human in the loop

we need a standaraizatin process, we need to trains, 所以为了design goal, 我们决定做好几层, design solution, 对应一条一条metric , so we need to make sure we achieve success metric 12345
每个solution怎么解决,
1. 一个section放一个success metric对应一个UI具体设计图, 文字结束
2. 一个section放一个success metric对应一个UI,

Problem area

For years, analysts collected data by hand — copying from source files and pasting into an internal tool. A single filing could take days.

Research surfaced two steps that ate the most time:

  1. Copying and pasting across documents

  2. Validating, comparing, and proofreading the result

We prioritized the second.
It saved the most time for the least effort. Automating the first would mean standardizing decades of inconsistent data formats across the industry — a far heavier lift for a smaller payoff.

The gaps compounded each other: no real-time validation, no error prevention, no way to track progress, constant context-switching, and no visual anchor to return to.


Design goal: How might we help analysts collect data faster — without sacrificing quality?

Technical Limitations

We had 30+ scattered collection tools, and adding AI to each one would cost too much. They had to become one tool with a shared schema first. I won leadership buy-in on that direction.

My framing shifted:
"We're not building an AI feature. We're building the foundation that makes AI possible — and the trust layer that lets researchers stop carrying the system's gaps with their own focus, memory, and time."

The Ideal

Full AI Automation

AI extracts all fields

Error Review

Done

GAP

The Reality: The fragmented landscape

No shared schema. No tracking. No quality check.

Excel

Survey

Scripts

Error List

Filing Viewer

Tracker

Intake Forms

Validation

Legacy Survey

Email

Survey

Manual Log

Vision: A unified platform where AI does the typing and humans do the judging.


Four design goals:
(1) Make AI extraction trustworthy at scale.
(2) Catch errors at entry, not downstream.
(3) Distinguish ownership of every flag.
(4) Design for two states of AI maturity at once.


Three design principles:
Error prevention is key.
Clarity over speed.
Keyboard friendly.
Human-in-the-loop all the time.

Solution overview - What Changed

From scattered tools to one source of truth. Errors started getting caught.

Before: scattered, 6 tabs, 15 hours, errors didn't get caught
Now: AI prefills, traceable sources, researchers review instead of type, catches errors before they propagate

yh92-1766834.github.io/prototype/
Click into the prototype — this is the launched running live.

THe Sequencing challenge

When our LLM team showed me that data accuracy varied significantly by datasets, I realized assuming AI confidence score as a single threshold made no sense.

That changed how I thought about the model itself. The AI was strong at structured field extraction from filings. It was weak at per-dataset accuracy variance, lack of historical context, and a tendency to be confidently wrong.

A hard decision: When should AI hand off to a human?

My starting hypothesis was simple: AI extracts most of the data, and analysts only validate the low-confidence cases.

But finance is different. The data types are rich and the accuracy bar is high — and testing showed researchers expected a different confidence threshold before the AI should hand off. So we defined the right moment to flag something for review.

Designing for a MOving TRAGET

The LLM team and I co-designed an AI maturity framework: three modes routed by confidence score — Edit, Co-pilot, Review — under one schema and one interaction model. My first sketches had only two modes. But confidence isn't static: it depends on the training data and improves as the model learns each dataset. So the modes are progressive — as accuracy stabilizes on a dataset, more of its data flows toward lighter-touch review. Researchers only ever see what genuinely needs them.

I weighed each error's impact, its risk level, and whether it could be fixed — because a rare error that corrupts a key value matters more than a common one that doesn't. My PM pushed for speed over quality; I pushed back with the analysis, and we landed on two error types: blocking errors that stop the workflow, and warning errors that flag without interrupting.


My impact: I built researcher trust in AI, earned strong stakeholder buy-in, and laid the foundation for both the AI data-collection platform and its design system.

Mode 1 - Edit Mode

Mode 2 - Review Mode

03 · Quality Check · Review Mode
Where Review Mode fits in the pipeline
01
Researcher Dashboard
Filing queue and personal stats. Pick up the next BDC filing and start a task.
02
Editor
AI-assisted data entry. Every value has a confidence score and a traceable source.
03
Quality Check
Review Mode — resolve flagged errors inline, with a side panel for the full queue.
04
Submission
Clean, validated data lands in the canonical dataset. One click to ship.
Optional flow
+
Borrower Resolution
Side flow when a reported borrower doesn't match any canonical PitchBook entity.

MODE 3 · Review Mode

Review Mode had one job: make the AI-to-human handoff legible — letting researchers resolve flags without losing their place in the queue.

I found that user would easily forget what was left, or lose context for the row.

A — Separate error panel

Cut

Value

!

Error

Field

Expected

Skip

Resolve

Why I cut it

Broke the researcher's position. Every resolution meant re-finding the row and remembering context.

B — Modal per error

Cut

Resolve error

×

VALUE

NOTE

Cancel

Save

Why I cut it

Dozens of flags per session. A modal per flag turned the queue into an interruption marathon.

C — Inline resolution

Cut

Subject

Value

VALUE

NOTE

Skip

Confirm

Why I cut it

Inline editing preserved row context but killed scan-mode — every flag collapsed the researcher's view of the whole queue into a single row. Fixing one error made her lose the other forty.

D — Side panel + inline

Shipped

Review Panel

Subject

Value

!

Why I shipped it

It gave researchers two surfaces working in parallel: scan-mode and edit-mode could coexist instead of replacing each other.

Side panel

Bottom drawer

Strategic decision

Enforced Quality Check before submission

There are actually two different quality check happening in the system: System-led and Human-led. The temptation was to skip the 2nd Quality check when a human edited the row. I pushed back. Errors don't only come from the AI. A researcher copying values across six tabs makes mistakes too.

So we enforced QC as a step, not assumed it as a property. The system treats every submission the same way.

Every rule I defined lives in the UI: pre-diagnosed errors, severity-ranked, click-to-jump, submission locked until clean.

The end-to-end data flow I mapped before designing a single screen — where AI extraction happens, where QC intercepts, where human judgment is required. This diagram didn't exist before I made it. It became the shared reference for design, engineering, and DataOps.

Trust by design: researchers can trace any AI value back to its source in one click. No black box.

What I navigated

The hardest work wasn't the system. It was the people around it.

A lot of the friction on this project came from unclear decision ownership. Early on, I wrote down what each function actually owned

Researchers → Fear of replacement

Researchers worried AI would replace them.

I showed them what the AI kept getting wrong. I made the case that their judgment, made systematic, was more valuable than their data entry. The researchers’ judgment was becoming the training signal for the system that would eventually reduce their manual work.

LLM team

Owned per-dataset accuracy and confidence thresholds.

Design

Owned the interaction model and trust layer.

PM

Owned scope and rollout sequencing.

What this changed

Naming this explicitly meant fewer arguments about whose call it was.

And made the disagreements that did happen substantive instead of territorial.

Design System

Impact

Shorter sessions, higher quality, and a platform to build on

For researchers
15+ hrs → minutes
Before
15+ hours per record across six tabs. No quality enforcement. Errors inherited downstream.
Now
Minutes per record on AI-handled data. One surface. Quality enforcement catches errors before they propagate.
For the product
0 → 1 unified platform
Before
No instrumentation, no shared schema, no path to AI extraction.
Now
The first unified data collection platform at PitchBook. MVP launched in Q1 2025 on the BDC dataset. Every future AI-powered tool builds on the schema and interaction model the team established.
For the design org
Reinvented → extended
Before
Designers reinvented patterns for each new collection tool.
Now
When three designers joined mid-project, they extended my system instead of starting over. The shared framework — table patterns, headers, navigation, entitlements — became the design language for every collection flow downstream.

FIntech