AudioSense
AudioSense — from WhatsApp voice notes to documents
A WhatsApp assistant that turns voice notes, text and images into transcriptions and Word or PDF documents.
- Project type
- In-house project
The project at a glance · Illustrated overview
From words to a document.

Messages
Voice notes, text and images on WhatsApp.

Transcription
Audio becomes usable text.

Context
Multiple inputs collected into one task.

Documents
Reports and notes in Word and PDF.

01
The challenge
Field professionals capture information as voice notes and photographs and need usable written documentation.
02
The solution
Webhooks receive messages and media, queued jobs process them, and OpenAI produces structured responses. Document-generation components assemble the output and deliver it through WhatsApp.
03
Capabilities
- Voice, text and image intake.
- OpenAI audio transcription.
- Structured responses and conversation context.
- Multiple contributions for reports, minutes and photo notes.
- Word and PDF generation and delivery.
- Message and webhook status tracking.
AUDIOSENSE / THE DIRECTION I’M BUILDING TOWARDS
From documents to operations
I want to bring this approach to future business applications, enabling staff to interact with systems and carry out complex operations through WhatsApp. This is the direction of my next projects, shaped around each client’s workflows.
LET’S BUILD WHAT’S NEXT
Start with the problem to solve.
A workflow to organise, systems to connect or a product to build. The first step is understanding the people, needs and constraints.