01The client
About ChapterMe
ChapterMe is an AI-native SaaS company based in Chicago. It helps content creators and podcasters work more productively by creating chapters for their YouTube videos and podcasts.
Well-written chapters make long-form content easier to navigate and easier to find, which increases content discovery, engagement and watch hours for every video they're added to.
02The system before
The system before: chapters written by hand
Chapters are the index of a long video. Producing good ones had been a slow, manual job:
Chapter creation
Writing chapters with descriptions and accurate timestamps took days.
Discovery and engagement
Without chapters, long videos and podcasts are harder to navigate and discover, which holds back engagement and watch hours.
Keeping pace with publishing
Creators who publish regularly can't spend days on chapters for every release.
A task that takes days per video can't keep up with a publishing schedule. It had to take minutes.
03The risk
Why it was risky to change
ChapterMe's customers publish what the model writes under their own names, to their own audiences. The output had to be right, not just fast.
RiskTimestamp accuracy
A chapter that starts at the wrong moment is worse than no chapter. Timestamps had to match the real change in content.
RiskLanguage quality
Chapters needed natural, descriptive titles and summaries a creator would put their name to, not keyword lists.
RiskYouTube compatibility
Chapters had to follow YouTube's format so they display correctly on the platform.
RiskAccount access
Creators link their YouTube accounts, so access had to stay limited to the creator's own videos.
04The architecture
The architecture: a multimodal model built for chapters
At the core is ChapterGPT, a custom-built multimodal model. For each segment of a video, it reads the transcript, detects visual elements in the video and identifies the emotion of the content, then combines those signals to decide where chapters begin and what they should say.
The output is a set of natural-language chapters with detailed descriptions and timestamps compatible with YouTube. Creators work through a web admin panel: they link their YouTube account and run ChapterGPT on the videos already in it.
What we engineered
-
ChapterGPT multimodal model
A custom-built model that reads transcripts, detects visual elements and identifies the emotion of each segment.
-
Natural-language chapters
Chapter titles with detailed descriptions, written the way a creator would write them.
-
YouTube-compatible timestamps
Every chapter starts at a timestamp formatted for YouTube.
-
Creator admin panel
A web panel where creators link their YouTube account and run ChapterGPT on the videos in it.
Key tradeoffWe built a multimodal model rather than chaptering from the transcript alone. Combining what is said, what is shown and how it feels takes more engineering, and it produces chapters that follow the content as viewers experience it.
05The results
The results
With chapters taking minutes rather than days, creators can add them to every video they publish, and every video can carry the discovery, engagement and watch-hour benefits chapters exist to deliver.
06Is this your situation?
Is your organization facing the same pattern?
This architecture applies to any product built on AI output. It fits organizations where:
- Your product depends on AI output that customers publish or act on under their own name.
- Video, audio, documents or transcripts are processed manually before they create value.
- Several signals, such as text, visuals and audio, need to be combined in one model.
- Your platform connects to customers' third-party accounts, and access has to be tightly scoped.
An architecture audit reviews your AI pipeline, integrations and data flows, and shows where engineering would move a measurable number.
07Questions
Frequently asked questions
What is ChapterGPT?
ChapterGPT is the custom-built multimodal model behind ChapterMe. It reads video transcripts, detects visual elements and identifies the emotion of each segment, then generates chapters with natural-language titles, detailed descriptions and YouTube-compatible timestamps.
Why use a multimodal model instead of the transcript alone?
A transcript only captures what is said. Adding what is shown on screen and the emotion of each segment lets the model place chapters where the content actually changes, and describe them more accurately.
How do creators use ChapterMe?
Through a web admin panel. Creators link their YouTube account and run ChapterGPT on the videos in it.
What results did the platform deliver?
Chapter creation went from days to minutes, ChapterMe's customer base grew by 210%, and the company was selected as one of the finalists for the Y Combinator Summer 2021 batch.
More case studies
-

Hospitality
Squarefare
Time to result cut by 98.4%; customer base up 42%.
Read the case study for Squarefare -

Edutech
Tektork
150 hours of manual work removed every month.
Read the case study for Tektork -

Fintech
Nagarathar Nalan
150% increase in donations.
Read the case study for Nagarathar Nalan