Design & media

声音克隆 Vidu Audio Clone

Clone a reference voice and generate audio that reads new text from the command line.

What it does

Clone a voice from reference audio and generate a reading of new text through the Vidu Audio Clone command. Provide an audio URL or local path with a prompt; the CLI uploads local media and sends the request to the hosted dLazy API. Results are returned as hosted output URLs, with optional asynchronous execution and status polling.

When to use it

  • Custom narration from reference audio
  • New script readings in a cloned voice
  • Asynchronous voice-generation jobs
  • CLI-based speech generation workflows

The skill document

声音克隆 Vidu Audio Clone

English · 中文

Clone voice and generate new text reading audio with one click using Vidu Audio Clone.

Trigger Keywords

  • vidu audio clone
  • clone voice
  • custom speech

Authentication

All requests require a dLazy API key. The recommended way to authenticate is:

dlazy login

This runs a device-code flow (also works in remote shells) and automatically saves your API key to the local CLI config — no manual copy/paste required.

Alternative: Set the Key Manually

If you already have an API key, you can save it directly:

dlazy auth set YOUR_API_KEY

The CLI saves the key in your user config directory (~/.dlazy/config.json on macOS/Linux, %USERPROFILE%\.dlazy\config.json on Windows), with file permissions restricted to your OS user account. You can also supply the key per-invocation via the DLAZY_API_KEY environment variable.

Getting Your API Key Manually

  1. Sign in or create an account at dlazy.com
  2. Go to dlazy.com/dashboard/organization/api-key
  3. Copy the key shown in the API Key section

Each key is scoped to your dLazy organization and can be rotated or revoked at any time from the same dashboard.

About & Provenance

You can install on demand without persisting a global binary by running:

npx @dlazy/cli@1.2.3 

Or, if you prefer a global install, the skill's metadata.clawdbot.install field declares the exact pinned version (npm install -g @dlazy/cli@1.2.3). Review the GitHub source before installing.

How It Works

This skill is a thin client over the dLazy hosted API. When you invoke it:

  • Prompts and parameters you provide are sent to the dLazy API endpoint (api.dlazy.com) for inference.
  • Any local file paths you pass to image / video / audio fields are uploaded to dLazy's media storage (files.dlazy.com) so the model can read them — the same flow as any cloud-based generation API.
  • Generated output URLs returned by the API are hosted on files.dlazy.com.

This is the standard SaaS pattern; the skill itself does not access network or filesystem resources beyond what the dLazy CLI already handles. See dlazy.com for the full service terms.

Usage

CRITICAL INSTRUCTION FOR AGENT: Run the dlazy vidu-audio-clone command to get results.

dlazy vidu-audio-clone -h

Options:
  --audio_url [audio_url]              Audio URL [audio: url or local path]
  --prompt [prompt]                    Prompt
  --dry-run                            Print payload + cost estimate without calling API
  --no-wait                            Return generateId immediately for async tasks
  --timeout                   Max seconds to wait for async completion (default: "1800")
  -h, --help                           display help for command

Any flag also accepts pipe references — - (auto-pick from upstream stdin), @N (n-th output), @N.path (jsonpath into output), @* (all primary values), @stdin / @stdin:path (whole envelope). See dlazy --help for details.

Output Format

{
  "ok": true,
  "result": {
    "tool": "vidu-audio-clone",
    "modelId": "vidu-audio-clone",
    "outputs": [
      {
        "type": "image",
        "id": "o_xxxxxxxx",
        "url": "https://files.dlazy.com/result.png",
        "mimeType": "image/png"
      }
    ]
  }
}

Async tasks (when --no-wait is passed) return outputs: [] and a task: { generateId, status } field instead. Use dlazy status --wait to poll.

Command Examples

# basic call:
dlazy vidu-audio-clone --prompt 'prompt text'

# complex call:
dlazy vidu-audio-clone --prompt 'prompt text'

Error Handling

CodeError TypeExample Message
401Unauthorized (No API Key)ok: false, code: "unauthorized", message: "API key is missing or invalid"
501Missing required parametererror: required option '--prompt ' not specified
502Local file read errorError: Image file/Video file not found: C:\path\to\your\file
503API request failed (no balance)ok: false, code: "insufficient_balance"
503API request failed (server error)HTTP status code error (500 server crash)
504Asynchronous task execution failed=== Generation Failed === / {Specific error reason returned by backend, for example "Prompt violates safety policy"}

AGENT CRITICAL INSTRUCTION:

  1. If the execution result returns code: "insufficient_balance", you MUST explicitly inform the user that their credits are insufficient and guide them to recharge: https://dlazy.com/dashboard/organization/settings?tab=credits
  2. If the execution result returns code: "unauthorized" or indicates missing API key, you MUST inform the user to get their API key from https://dlazy.com/dashboard/organization/api-key and save it using dlazy auth set and resume the task.

Tips

Visit https://dlazy.com for more information.

Questions people ask

What inputs does voice cloning accept?
The command accepts reference audio as either a URL or a local file path through `--audio_url`, plus the text or instructions through `--prompt`.
Do I need an API key?
Yes. All requests require a dLazy API key, which can be saved with `dlazy login` or `dlazy auth set YOUR_API_KEY`, or supplied through `DLAZY_API_KEY`.
Can I run the generation asynchronously?
Yes. Use `--no-wait` to receive a generation ID immediately, then poll it with `dlazy status --wait`. The default synchronous timeout is 1,800 seconds and can be changed with `--timeout`.

Related skills

Generate Vidu Q2 images from prompts and references at configurable aspect ratios and resolutions.

63 installs

Convert Chinese or English text into emotional speech with selectable voices, speeds, and MP3 or WAV output.

65 installs

Generate 3–15 second Kling V3 videos from text or images with configurable format, mode, and sound.

65 installs

Turn up to 10 images or selected first and last frames into a Vidu Q2 video with configurable output settings.

61 installs

Generate narration, cloned voices, sound effects, and music from text prompts.

225 installs4 stars