Web scrapers

Transcript API

One transcript shape across every platform that has one. Each endpoint returns the spoken content of a video or an audio post as timestamped segments, so a TikTok video, a YouTube upload, an Instagram reel, a Facebook post and an X video all arrive as the same array of lines with a start and an end time. That is the step that makes video searchable. Once the words are text you can index them, translate them, hand them to a model, or find the exact second a phrase was said.
  • 5 endpoints
  • 5 platforms
  • from $1.49 per 1,000 calls
  • One key, one JSON shape

In every response

The same JSON, whichever platform you ask.

  • segments[] Each transcript line with its start/end timestamp and text.
  • language The detected or declared transcript language.
  • duration Total media length the transcript spans.
  • source Whether the transcript is creator-provided or machine-generated.
curl "https://hproxy.com/api/v1/scrape/facebook/post-transcript?url=https%3A%2F%2Fwww.facebook.com%2Fnasa" \
  -H "X-API-Key: hpx_your_key_here"
Response · Facebook Post Transcript200 OK · JSON
{
  "language": "en",
  "duration": 47.2,
  "source": "creator",
  "segments": [
    { "start": 0.0, "end": 2.4, "text": "So the first thing we changed was the timing," },
    { "start": 2.4, "end": 5.8, "text": "and that completely changed the result." }
  ]
}

What people build with it

Where transcript api earns its keep.

  • Index a creator's whole catalogue so the words spoken inside videos become searchable text.
  • Feed clean transcripts to a summarisation or classification model with no speech-to-text bill.
  • Find the exact timestamp where a product, a competitor or a claim is mentioned.
  • Generate subtitles or translations from the creator's own words rather than a re-recording.

Transcript API questions

Answered before you ask.

The scraper docs

No. You send the post URL and the transcript comes back as JSON. The media never touches your infrastructure and there is no speech-to-text step or bill on your side.

The source field tells you which, so you can decide how much to trust it. Creator-provided caption files are near-perfect; automatic captions are good enough for search and rough for quoting.

Not every post carries one. Where the platform holds no transcript for that media, the call reports that rather than returning invented text, and an extraction that cannot complete is not charged.

Yes. Every segment carries its own start and end, which is what you need to seek to a phrase or slice a clip around it, rather than a single wall of text with no positions.

More on transcript api

What transcripts covers

5 endpoints across 5 platforms: TikTok Video Transcript, Facebook Post Transcript, YouTube Video Transcript, Instagram Media Transcript and Twitter / X Tweet Transcript. They are separate scrapers behind one API key, and they answer in the same JSON shape, so the code that reads one reads all of them.

Why the shape matters

The spoken words inside video, as timestamped JSON lines. Working platform by platform means a second integration for every network you add. Here the field names, the pagination and the error codes are the same wherever the data came from, so adding another platform is a change of URL rather than a rewrite.

What it costs

Calls in this group start at $1.49 per 1,000 and are billed from one balance, per call, with no subscription. Blocked, errored and timed-out calls are not charged, and each endpoint prices itself: a heavier lookup costs more than a simple one, and the page for each endpoint prints its own rate.

Ready when you are.

Your dashboard is ten seconds away. No sales call, no subscription, no minimum.