← All posts

Tutorial

AI Video API Webhooks: Design, Security, and Recovery

A completion callback is only the beginning of delivery. Learn how to verify video webhook events, store them durably, handle duplicates, recover missed notifications, and turn completed generations into usable assets.

Last updated: September 28, 2026

A video generation task finishes while your application is deploying. The provider sends a callback, but your receiver is temporarily unavailable.

In another request, the callback reaches your server successfully. Your application downloads the video, then crashes before recording that processing is complete. The provider retries, and the same completion event arrives again.

Both situations are normal failure cases for asynchronous systems. Neither should leave a customer waiting indefinitely or cause your application to create another video unnecessarily.

AI video API webhooks notify your backend about an existing generation task. A reliable implementation checks whether the notification can be trusted, records it durably, and processes the result through a recoverable workflow.

The essential sequence is receive, verify, persist, acknowledge, process, and reconcile.

This guide explains how to design that sequence, handle repeated and late events, and recover when no callback arrives. It also separates general webhook practices from the specific behavior documented by individual providers.

AI Video API Webhooks: Design, Security, and Recovery — workflow illustration

Before You Read

This article assumes that your application can submit a video request and retain its task ID.

If you are still building that foundation, start with:

If your application already maintains durable video jobs, continue below to connect provider notifications to that system.

AI Video API Webhooks at a Glance

These stages do not require six services. A small application can use one receiver, a database inbox, a background worker, and a scheduled recovery process.

The important boundary is between accepting a notification and completing the business work it triggers.

What an AI Video Webhook Actually Does

A webhook is a separate HTTP request from the provider to an endpoint controlled by your application. It reports information about work that was submitted earlier.

It does not replace your job database, establish user permissions, or guarantee that a finished video is ready to display.

For example, fal documents callbacks for asynchronous queue requests. Developers supply a callback destination, and fal posts the result when processing completes. Its webhook documentation defines the payload and verification procedure for that integration.

That is one provider contract. Its fields and authentication mechanism should not be copied into another API.

Keep two milestones separate:

Generation completed: The provider has finished producing the output.

Application ready: Your system has retrieved the output, completed required checks, and made it accessible to the correct user.

A callback can trigger the transition between them. It should not collapse them into a single assumption.

When Should You Use Polling or Webhooks?

Polling and webhooks provide different ways to discover task state.

Polling asks the provider whether an existing task has changed. Webhooks allow the provider to initiate a notification.

For an initial integration, a polling worker may be sufficient. Add webhooks when they reduce unnecessary checks or make background completion handling easier.

Webhooks do not make the model generate faster. They change how your application learns that the work has finished.

For production systems, the useful question is whether the notification mechanism has a recovery path when delivery fails.

Define the Callback Contract Before Writing the Handler

Primary purpose: Establish exactly what your receiver can rely on.

Start with the documentation for the specific endpoint. Record:

  • How callbacks are registered.

  • Which events are emitted.

  • How the sender is authenticated.

  • Which identifier remains stable across retries.

  • What response acknowledges delivery.

  • How long the provider waits and retries.

  • How task state and outputs can be recovered afterward.

Distinguish the identifiers involved.

A task ID identifies generation work. An event ID identifies a logical notification. A delivery attempt ID may identify only one HTTP attempt.

If every retry receives a new delivery ID, deduplicating on that value will not prevent repeated processing of the same event.

Your internal event record should include the provider account, task ID, stable event key where available, event type, receive time, trust state, and processing state.

Associate the task with its workspace through your own submission records. A user or workspace identifier inside a callback must not independently grant access.

Implementation consideration: Retain only the payload information needed for processing and investigation. Signed asset URLs and sensitive input references should not appear in ordinary logs.

Verify the Notification Before Trusting Its Contents

Primary purpose: Prevent external requests from directly controlling application state.

Where the provider signs deliveries, preserve the raw request bytes and follow its documented verification procedure.

Parsing JSON and then serializing it again can change the bytes covered by a signature. A verifier should use the required original representation.

GitHub’s documentation illustrates HMAC validation and constant-time comparison in its guidance on validating webhook deliveries. The principle is useful, but GitHub’s header names and signing scheme are specific to GitHub.

Where a signed timestamp is part of the protocol, validate its age within the provider’s documented rules. An unsigned timestamp does not prove freshness.

Signature verification and deduplication solve different problems. An authentic event can still arrive more than once.

If the endpoint has no authenticated callback mechanism, treat the notification as an untrusted hint. Restrict it to known tasks, apply request limits, and verify the current task through an authenticated API read before changing trusted job state or downloading an output.

Implementation consideration: HTTPS protects transport, but it does not by itself prove that an incoming request came from the provider.

Persist the Event Before Acknowledging Delivery

Primary purpose: Make accepted notifications recoverable after a process failure.

Return a successful acknowledgment only after the event has been durably accepted.

Video downloads, media inspection, and user notifications belong in a background worker. Performing them inside the callback request increases acknowledgment time and connects delivery success to unrelated downstream failures.

Stripe documents quick acknowledgment, duplicate events, and ordering considerations in its webhook handling guidance. These are useful design references; its retry schedule is not a specification for video providers.

A practical pattern is a database inbox: a table containing accepted events and their processing state. Workers claim pending records using recoverable leases.

If a worker crashes, another can resume after the lease expires.

For a signed callback, the receiver boundary can be represented as follows:

receive(request):
    enforce_request_limits()

verified = adapter.verify(
        request.raw_body,
        request.headers
    )

event = adapter.normalize(verified)

begin transaction
        insert event if its stable key is new
        preserve existing work if it is a duplicate
    commit

return documented_success_response

This is application pseudocode, not a runnable provider endpoint.

For unsigned callbacks, use a separate path that stores a bounded, untrusted hint associated with a known task. A worker then confirms status through the provider API before creating trusted completion work.

Implementation consideration: Saving an event and separately sending a queue message creates a failure gap. Let workers scan the inbox, or save an outbox record in the same transaction.

Make Completion Processing Safe to Repeat

Primary purpose: Prevent repeated notifications or worker retries from producing duplicate effects.

Receiver deduplication is only the first layer.

A worker can save a video successfully and crash before marking the event processed. When processing resumes, the application must recognize the completed storage operation.

Use an asset record or deterministic storage key tied to the internal job and generation attempt. Protect the transition to ready with a conditional update.

Handle notifications separately. Record a distinct notification operation and use downstream idempotency support where available. A local record alone cannot guarantee exactly-once delivery through an external service.

Do not calculate billable usage from callback counts. Repeated callbacks are transport behavior, not additional model execution.

When no stable event ID exists, define deduplication from documented semantics. A task ID may identify one completion operation, but it can be too broad for an endpoint that emits several meaningful events for the same task.

Implementation consideration: Keep event deduplication and business-operation deduplication separate. Two different events may legitimately refer to the same completion work.

Handle Late Events Without Reversing a Completed Job

Primary purpose: Keep delayed updates from corrupting current application state.

Receipt order does not necessarily establish event order.

Use explicit transition rules instead of assigning whichever status arrived most recently.

Keep each generation attempt distinct.

If a user has requested a replacement, a late output from the earlier attempt should not automatically overwrite the selected result or trigger an unexpected notification.

Cancellation requires similar care. A cancelled interface state may mean the user stopped waiting, even if the provider continued executing.

Implementation consideration: Quarantine unmatched or conflicting events rather than guessing ownership or selecting an output based only on arrival time.

Recover When No Webhook Arrives

Primary purpose: Prevent missed delivery from leaving completed jobs permanently unresolved.

Run a reconciliation worker for jobs whose next status check is due.

It should inspect existing tasks, recover missed state changes, and resume unfinished asset handling.

Schedule checks around observed generation times, provider limits, and output retention. Use bounded backoff and jitter for temporary read failures.

A user-facing deadline may expire while upstream generation is still running. Preserve that uncertainty rather than converting it into a confirmed failure.

When reconciliation finds a completed task, send it through the same completion operation used by callbacks. Both paths should converge on the same asset and state transition.

Token360’s rate limits, timeouts, and retries guidance distinguishes temporary capacity conditions from quota or spending restrictions and warns against blind resubmission after ambiguous timeouts.

Implementation consideration: A missing callback alone is not a reason to generate another video. Investigate the existing task first.

Finish Asset Delivery Before Showing Ready

Primary purpose: Turn an upstream completion into a usable customer result.

Retrieve the output from a trusted or independently verified task result.

Apply download timeouts and size limits, validate the media, and store the asset under the correct workspace. If externally influenced URLs are involved, prevent access to private network destinations and validate redirects.

Use the documented retrieval mechanism. Token360 provides a dedicated video download reference for its download endpoint.

Record the storage location and relevant metadata before announcing readiness. If downloading fails, retry retrieval while the output remains available.

Keep that failure distinct from generation failure. Otherwise, a temporary storage problem can accidentally create another paid generation.

Where moderation or review is required, ready should reflect those checks too. A completion callback does not establish creative quality or publication approval.

Implementation consideration: Define retention for callback payloads, downloaded files, and logs separately. Token360’s security and data handling guidance identifies these as distinct parts of the data path.

Test the Failures That Leave Customers Waiting

Primary purpose: Verify recovery without depending on repeated paid generation.

Use synthetic events and controlled failures to test the receiver and worker independently.

The following is a proposed validation plan, not a report of measured platform results.

Measure acknowledgment latency, oldest pending event age, worker retry counts, reconciliation recoveries, and time from upstream completion to application readiness.

A receiver can return fast responses while its processing queue is stalled. Monitor both stages.

Implementation consideration: Use separate alerts for failed delivery, failed processing, and unresolved task state. They require different interventions.

How Token360 Video Webhooks Fit This Design

Token360’s current video webhook documentation describes callback_url on POST /v1/videos and an optional business request_id echoed in the callback.

Video callbacks report completed or failed, rather than intermediate progress. The documentation specifies up to three delivery attempts, a 30-second timeout per attempt, and acknowledgment through a successful 2xx response.

The video callback section states that deliveries carry no authentication headers. The same page documents optional signing for Batch callbacks; that mechanism should not be assumed to apply to video.

For video integration, treat callbacks as hints tied to known tasks and verify status through the authenticated task endpoint before committing consequential state changes.

Also, reusing request_id does not deduplicate multiple generation tasks. Keep task correlation, event deduplication, and submission idempotency as separate concerns.

Which Implementation Should You Start With?

If your application already polls reliably, keep that recovery path and add callbacks as a way to discover completion sooner.

For a small system, a receiver, durable inbox, background worker, and scheduled reconciler are often sufficient. Add dedicated queues or separate processing services when workload measurements justify them.

For multiple providers, share the internal completion workflow while keeping registration, verification, payload normalization, and status mapping inside provider-specific adapters.

This preserves a consistent application experience without pretending that every webhook contract is identical.

The implementation is ready for production when a repeated callback, missed callback, or worker crash leads to recoverable work rather than a stranded job.

Frequently Asked Questions About AI Video API Webhooks

Are webhooks better than polling for video generation?

They serve different needs. Polling is straightforward when a status endpoint exists. Webhooks can reduce routine checks but require a receiver with durable processing and recovery.

Should I download the video inside the webhook request?

Usually, accept the notification durably and download through a worker. This keeps acknowledgment fast and allows retrieval failures to recover independently.

Does signature verification prevent duplicate processing?

No. It checks authenticity and integrity under the signing protocol. An authentic event can still be delivered repeatedly.

What if every retry has a different delivery ID?

Use a stable event identity or a business-operation key based on documented semantics. A changing delivery ID cannot identify duplicates by itself.

What if the callback has no signature?

Treat it as an untrusted hint. Match it to a known task and verify the task through an authenticated API request before applying trusted updates.

Can one handler support every video provider?

You can share durable storage and completion processing. Keep provider-specific authentication, normalization, and event semantics separate.

What if the endpoint does not support webhooks?

Use a background polling worker. After verifying completion, it can invoke the same asset-processing operation that a callback would trigger.

What to Read Next

Once callback receipt and recovery are reliable, continue with the parts of the workflow that determine what happens after a failure or across many tasks:

  • AI Video API Error Handling — Blog 15: Define how to distinguish invalid requests, temporary failures, uncertain submissions, and deliberate generation retries.

  • Batch Video Generation Guide for Partial Completion and Recovery: Extend recovery to groups of tasks without repeating items that already succeeded.

  • Rate Limits, Timeouts, and Retries: Review the documented operational conditions before implementing retry rules.

The next step is to connect notification handling with a consistent policy for task failures, output recovery, and workload limits.

Ready to Connect Video Callbacks to Your Application?

Start with the documented video submission workflow, then connect callbacks and reconciliation to the same durable job record.

Verify the result, preserve the asset, and mark the job ready only when your application can deliver it.

Review the Token360 video API workflow →


  • Video Generation
  • API
  • Developer Guide

Build faster with one AI API.

Use Token360 to call video, image, audio, and text models with one key and one bill.

Get started