Metadata-Version: 2.4
Name: api-sdk
Version: 0.2.0
Summary: API SDK for image generation and other services
Author-email: Your Name <your.email@example.com>
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: elevenlabs>=1.0.0
Requires-Dist: google-genai>=1.56.0
Requires-Dist: numpy>=2.0.0
Requires-Dist: onnxruntime>=1.20.0
Requires-Dist: opencv-python>=4.10.0
Requires-Dist: Pillow>=10.0.0
Requires-Dist: pymatting>=1.1.13
Requires-Dist: rembg>=2.0.60
Requires-Dist: requests>=2.31.0
Requires-Dist: volcengine-python-sdk[ark]>=5.0.42
Requires-Dist: websocket-client>=1.6.0

# API SDK

A Python SDK for image generation using Google Gemini API.

## Installation

```bash
pip install -e .
```

Or install dependencies directly:

```bash
pip install -r requirements.txt
```

## Configuration

This SDK does not load any `.env` file. Provide credentials via environment variables:

```bash
# Google Gemini
export GOOGLE_API_KEY=your-api-key
# (Optional legacy alias) export GEMINI_API_KEY=your-api-key

# Volcengine Jimeng
export VOLC_ACCESSKEY=...
export VOLC_SECRETKEY=...

# Volcengine Ark Seedream 5.0 Lite
export ARK_API_KEY=...

# TTS providers
export ELEVENLABS_API_KEY=...
export VOLCENGINE_TTS_API_KEY=...
export VOLCENGINE_TTS_APP_ID=...  # optional; set it when your console credentials include one
export FLYTEK_APPID=...
export FLYTEK_APIKey=...
export FLYTEK_APISecret=...
export FLYTEK_URL=wss://tts-api-sg.xf-yun.com/v2/tts
```

## Usage

### As a Python Library

```python
from api_sdk.image_generator import generate_image
from api_sdk import generate_text
from api_sdk import generate_speech, generate_speech_bytes

# Generate an image
output_path = generate_image(
    prompt="A beautiful sunset over mountains",
    output_path="./images/sunset.png"
)
print(f"Image saved to: {output_path}")

# Generate text
text, _ = generate_text(
    prompt="给我一首关于秋天的短诗",
)
print(text)

# Synthesize speech (bytes, no disk I/O)
audio_bytes = generate_speech("你好，世界")  # 默认返回 WAV bytes
Path("./audio").mkdir(parents=True, exist_ok=True)
Path("./audio/hello.wav").write_bytes(audio_bytes)
```

### Using the CLI

Generate an image with a simple prompt:

```bash
python -m api_sdk.cli generate-image "A beautiful sunset over mountains"
```

Generate an image and save to specific location:

```bash
python -m api_sdk.cli generate-image "A cat playing piano" -o ./images/cat_piano.png
```

Use verbose mode for detailed output:

```bash
python -m api_sdk.cli generate-image "Modern architecture" -v
```

No CLI `--api-key` flag is provided; credentials are discovered automatically
via environment variables.

After installation with pip, you can also use the command directly:

```bash
# Default image provider is Jimeng (requires VOLC_ACCESSKEY/VOLC_SECRETKEY)
api-sdk generate-image "Your prompt here"

# Or use Gemini explicitly
api-sdk generate-image --provider gemini "Your prompt here"
```

Generate image (Volcengine Jimeng / 即梦，AK/SK):

```bash
export VOLC_ACCESSKEY=...
export VOLC_SECRETKEY=...

# 文生图（当前接入：jimeng_t2i_v40）
api-sdk generate-image --provider jimeng "公共信息墙上的医学文献检索结果页面，冷静权威可信" -o ./images -n jimeng_img

# 可选：尝试指定宽高（是否支持取决于模型；高级参数可用 --extra-config-file 透传）
api-sdk generate-image --provider jimeng "A calm authoritative information dashboard UI" --width 1024 --height 576 -o ./images -n jimeng_16_9
```

Generate image with Seedream 5.0 Lite through Volcengine Ark:

```bash
export ARK_API_KEY=...

# Text-to-image
api-sdk generate-image --provider seedream \
  "1930s mountain road after rain, restrained historical documentary style" \
  --width 2560 --height 1440 -o ./images -n seedream_landscape

# Single or multiple local reference images (up to 14)
api-sdk generate-image --provider seedream \
  "Keep the clothing and landscape style of the references; create a new wide shot" \
  -i ./references/style.png -i ./references/character.png \
  --width 2560 --height 1440 -o ./images -n seedream_referenced
```

Generate video (Gemini Veo):

```bash
api-sdk generate-video "A cinematic shot of a tiny robot walking through a rainy neon city"
api-sdk generate-video "Make it sunrise" -i examples/1.jpg -o ./videos -n sunrise

# With reference images (max 3) + optional types (ASSET / STYLE)
api-sdk generate-video "Keep the same character, new scene" \
  --reference-image examples/1.jpg --reference-type ASSET \
  --reference-image examples/2.jpg --reference-type STYLE \
  -o ./videos -n ref_demo

# Extend an existing video
api-sdk generate-video "Continue the scene with slower camera movement" \
  --extend-video ./videos/demo.mp4 \
  --duration-seconds 8 --resolution 720p -o ./videos -n extended
```

Generate video (Volcengine Jimeng / 即梦，AK/SK):

```bash
# 通过环境变量提供火山引擎 AK/SK
export VOLC_ACCESSKEY=...
export VOLC_SECRETKEY=...

# 文生视频（当前接入：jimeng_t2v_v30）
api-sdk generate-video --provider jimeng "一段冷静、权威的医学文献检索信息墙界面，从加载态过渡到结果总览" \
  --aspect-ratio 16:9 --duration-seconds 5 -o ./videos -n jimeng_demo

# Jimeng 也可直接指定帧数（121=5s，241=10s）
api-sdk generate-video --provider jimeng "A calm authoritative public information panel UI" --frames 241 -o ./videos -n jimeng_10s
```

Generate video (Volcengine Ark Seedance，API Key):

```bash
export ARK_API_KEY=...

# 文生视频；默认使用 Seedance 2.0 标准版
api-sdk generate-video --provider seedance \
  "1930年代山路，克制的历史纪录片质感，镜头缓慢前移" \
  --aspect-ratio 16:9 --resolution 720p --duration-seconds 5 \
  -o ./videos -n seedance_demo

# 首帧/尾帧控制；Fast 模型可显式指定
api-sdk generate-video --provider ark \
  "保持人物与环境一致，只做自然微动和缓慢运镜" \
  --model doubao-seedance-2-0-fast-260128 \
  --image ./frames/start.png --last-frame ./frames/end.png \
  --camera-fixed --duration-seconds 5 \
  -o ./videos -n seedance_keyframes
```

Python 调用使用相同参数；Ark 特有开关通过 `extra_config` 传入：

```python
from api_sdk import generate_video

video_path, description_path = generate_video(
    "subtle documentary motion",
    provider="seedance",
    input_image="./frames/start.png",
    duration_seconds=5,
    extra_config={
        "camera_fixed": True,
        "watermark": False,
        "return_last_frame": False,
    },
)
```

Seedance 当前每个调用生成一个视频，时长支持 4–15 秒。尾帧必须与首帧同时提供。
`reference_images`、视频续写和 Veo 专属配置会直接报错，不会静默忽略。

Generate text:

```bash
api-sdk generate-text "给我一段关于海风的散文"
api-sdk --json generate-text "Summarize: ..." --model gemini-2.5-pro --max-tokens 512

# Use Codex as text provider (non-interactive, safe sandbox)
api-sdk generate-text "写一首关于红豆的诗歌（中文）" --provider codex
python -m api_sdk.cli generate-text --provider codex "写一句关于红豆的短句（中文）"
```

Structured Output (Gemini):

```bash
# Pass a JSON schema to enforce structured JSON output.
api-sdk generate-text "Generate a recipe for butter cookies" --provider gemini --response-schema-file ./examples/recipe.schema.json
```

Structured Output (Codex):

```bash
api-sdk generate-text "Extract project metadata" --provider codex --response-schema-file ./examples/recipe.schema.json
```

Models:
- Default image provider: `jimeng`
- Default image model (Jimeng req_key): `jimeng_t2i_v40`
- Seedream/Ark image model: `doubao-seedream-5-0-260128`
- Gemini image default model: `gemini-3.1-flash-image-preview` (when `--provider gemini`)
- You can also pass a different model explicitly, e.g.: `--model gemini-2.5-flash-image` or `--model jimeng_t2i_v40`
- Default video model (Gemini Veo): `veo-3.1-generate-preview`
- Default video model (Jimeng req_key): `jimeng_t2v_v30`
- Default video model (Ark Seedance): `doubao-seedance-2-0-260128`
- Ark Seedance Fast model: `doubao-seedance-2-0-fast-260128`

Global flags:
- `--json` Emit machine-readable JSON to stdout (for pipelines)
- `-v/--verbose` Verbose logs (DEBUG)
- `-q/--quiet` Suppress non-error logs
- `--no-color` Disable decorative icons/emojis

Example JSON output:

```bash
api-sdk --json detect-objects examples/1.jpg "Find animals"
# {"ok":true,"command":"detect-objects","output_path":"images/1_detected.jpg","objects":[...],"model":"gemini-2.0-flash-exp"}
```

Shell wrapper:
- A thin shell launcher is available at `bin/api-sdk` (executes `python -m api_sdk.cli "$@"`).
  You can run `chmod +x bin/api-sdk` and add it to your PATH if desired.

### Text-to-Speech (TTS)

支持提供方：Google Gemini TTS（默认，Algieba）、火山引擎豆包语音、iFlytek、ElevenLabs。

火山引擎豆包语音 V3（默认使用 2.0 音色 `zh_female_vv_uranus_bigtts`，24 kHz）：

```bash
export VOLCENGINE_TTS_API_KEY=...
# 若当前控制台凭据包含 APP ID，则一并设置；SDK 会自动发送。
export VOLCENGINE_TTS_APP_ID=...

api-sdk synthesize-speech "先完成最小的一步，我们再决定下一步。" \
  --provider volcengine --format mp3 -o ./audio/coach.mp3
```

可通过 `--voice` 或 `VOLCENGINE_TTS_VOICE` 更换音色。高级配置可使用：

- `VOLCENGINE_TTS_RESOURCE_ID`：默认 `seed-tts-2.0`；
- `VOLCENGINE_TTS_ENDPOINT`：默认 V3 SSE 单向流式端点；
- `VOLCENGINE_BIDIRECTIONAL_TTS_ENDPOINT`：默认 V3 双向流式 WebSocket 端点；
- `VOLCENGINE_TTS_UID`：可选的非敏感业务用户标识；
- `--speed`：语速 `-50...100`；
- `--volume`：响度 `-50...100`；
- `--pitch`：音调 `-12...12`。

需要边接收边播放时，可直接消费 V3 SSE 返回的 24 kHz、16-bit、单声道 PCM 分片；调用 `close()` 可停止并释放当前响应：

```python
from api_sdk import create_volcengine_streaming_tts_session

session = create_volcengine_streaming_tts_session(text="这是最终答复。")
try:
    for pcm_s16le_chunk in session.audio_chunks():
        play_or_forward(pcm_s16le_chunk)
finally:
    session.close()
```

`audio_chunks()` 只能消费一次，并要求服务端发出成功完成帧；完整 WAV/MP3 调用仍使用 `synthesize_speech_volcengine_bytes()`。

LLM 边生成文字、边合成语音时使用 V3 双向会话。一个逻辑答复只建立一次 WebSocket，后续稳定文本段通过 `append_text()` 追加到同一 session；`finish()` 表示文本输入结束，`close()` 用于取消：

```python
from api_sdk import create_volcengine_bidirectional_tts_session

session = create_volcengine_bidirectional_tts_session(uid="my-agent-chat")
try:
    session.append_text("第一段。")
    session.append_text("第二段仍在生成。")
    session.finish()
    for pcm_s16le_chunk in session.audio_chunks():
        play_or_forward(pcm_s16le_chunk)
finally:
    session.close()
```

双向会话默认连接 `wss://openspeech.bytedance.com/api/v3/tts/bidirection`，`provider_log_id` 暴露握手返回的 `X-Tt-Logid` 供无正文诊断。单向 SSE 要求一次性输入完整文本，不能按句段重复建连来实现 LLM 实时朗读。接口选择以火山官方 [语音合成大模型 API 列表](https://www.volcengine.com/docs/6561/2228192?lang=zh) 和 [V3 双向流式文档](https://www.volcengine.com/docs/6561/2532486?lang=zh) 为准。

ElevenLabs（需要在环境变量里配置 `ELEVENLABS_API_KEY`；voice_id 默认 `DowyQ68vDpgFYdWVGjc3`，也可通过 `ELEVENLABS_VOICE_ID` 或 CLI `--voice` 覆盖）：

```bash
export ELEVENLABS_API_KEY=...
export ELEVENLABS_VOICE_ID=JBFqnCBsd6RMkjVDRZzb

api-sdk synthesize-speech "The first move is what sets everything in motion." \
  --provider elevenlabs --format mp3 --model eleven_multilingual_v2 -o ./audio/eleven.mp3
```

合成中文语音示例（默认输出 WAV 16k 单声道）：

```bash
# Gemini（默认，Algieba）
api-sdk synthesize-speech "你好，世界" -o ./audio/hello_google.wav --format mp3

# 指定模型（默认已是 gemini-2.5-pro-preview-tts）：
api-sdk synthesize-speech "你好，世界" --provider gemini --model gemini-2.5-pro-preview-tts -o ./audio/hello2.mp3 --format mp3

# 生成 MP3（Google 侧支持）：
api-sdk synthesize-speech "Hello, this is a demo." --provider algieba --model gemini-2.5-pro-tts --format mp3 -o ./audio/demo.mp3
```

说明：
- Google Gemini TTS（Algieba）默认输出可直接为 MP3（--format mp3）。可选传入 `--voice`（可用值因模型/地区不同）。
- iFlytek 默认音色（voice/vcn）为 `x_xiaoyang_story`，也可以通过 `--voice` 自行指定。

自定义参数（音色、格式、采样率）：

```bash
api-sdk synthesize-speech "欢迎使用语音合成" \
  --voice x_John \
  --format wav \
  --sample-rate 16000 \
  -o ./audio/welcome.wav
```

环境变量:
- `FLYTEK_APPID`
- `FLYTEK_APIKey`
- `FLYTEK_APISecret`
- `FLYTEK_URL`（可选，覆盖默认端点）
- `VOLCENGINE_TTS_API_KEY`
- `VOLCENGINE_TTS_APP_ID`（可选）
- `VOLCENGINE_TTS_VOICE`（可选）
- `VOLCENGINE_TTS_RESOURCE_ID`（可选）
- `VOLCENGINE_TTS_ENDPOINT`（可选）
- `VOLCENGINE_BIDIRECTIONAL_TTS_ENDPOINT`（可选）
- `GEMINI_API_KEY`（可选；若未设置则走 Google 默认凭据/ADC）

认证说明（Google）：
- 本地开发可直接设置 `GEMINI_API_KEY`（无需交互），或运行 `gcloud auth application-default login` 使用 ADC（需要一次交互登录）。
- 生产环境推荐使用服务账号，通过 `GOOGLE_APPLICATION_CREDENTIALS` 指向 JSON key，或使用云环境的工作负载身份（无人工交互）。

Provider 别名：`google`、`gemini`、`algieba` 均指向 Google Gemini TTS；`volcengine`、`doubao` 均指向火山引擎豆包语音。默认 provider 为 `algieba`。

注意：可用的音色代码（`--voice`/vcn）因账号而异，请根据你在讯飞控制台开通的音色填写。

### Speech-to-Text (ASR)

火山 Seed-ASR 2.0 优化双向流式接口支持实时发送 PCM 并接收不断修订的识别全文。默认使用小时版资源 `volc.seedasr.sauc.duration`：

```python
from api_sdk import create_volcengine_streaming_asr_session

session = create_volcengine_streaming_asr_session(
    api_key="...",  # 也可使用下述环境变量
    sample_rate=16000,
    channels=1,
    bits_per_sample=16,
)
try:
    for result in session.send_audio(pcm_s16le_chunk):
        print(result.text, result.is_final)
    final_results = session.finish()
finally:
    session.close()
```

每个音频 chunk 必须是 16 kHz、16-bit、单声道 little-endian PCM；`send_audio()` 返回当前独立接收循环已经取得的结果，`text` 是全文 replacement，不应直接追加，尚未到达的结果会由后续 `send_audio()` 或 `finish()` 收口。默认请求开启 `enable_ddc=true` 语义顺滑和 `enable_nonstream=true` 二遍识别，默认 endpoint 为 `wss://openspeech.bytedance.com/api/v3/sauc/bigmodel_async`。会话的 `provider_log_id` 暴露握手响应 `X-Tt-Logid` 供日志关联，但不保存音频或转写正文。可选配置：

- `VOLCENGINE_STREAMING_ASR_RESOURCE_ID`：默认 `volc.seedasr.sauc.duration`；
- `VOLCENGINE_STREAMING_ASR_ENDPOINT`：覆盖流式 WebSocket endpoint。

协议参数与分包要求以火山引擎官方文档为准：
[双向流式语音识别 WebSocket](https://www.volcengine.com/docs/6561/2630027?lang=zh)、
[大模型流式语音识别 API](https://www.volcengine.com/docs/6561/1354869?lang=zh)。

## API Reference

### `generate_image()`

Generate an image based on a text prompt.

**Parameters:**
- `prompt` (str): Text prompt for image generation
- `output_path` (str, optional): Path where the generated image will be saved
- `provider` (str, optional): Image provider (`jimeng` by default, or `gemini`)
- `api_key` (str, optional): Google Gemini API key (Gemini only)
- `model` (str, optional): Model to use for generation

**Returns:**
- `(image_path, description_path)`: Paths to the saved image and optional description markdown

### `ImageGenerator` Class

A wrapper class for image generation (Gemini or Volcengine Jimeng / 即梦).

```python
from api_sdk.image_generator import ImageGenerator

generator = ImageGenerator(api_key="your-api-key", provider="gemini")
image_path, description_path = generator.generate_image(
    prompt="Your prompt",
    output_path="output.png"
)
```
