Go to file

Rdzleo b8a5fe958f feat(rtc-only): Phase 6 - RTC 空闲软休眠（B+C 双源 + 真退房 + 字幕提示 + 内存兜底）

按 GSD 框架 .planning/milestones/digital_human_rtc/phases/phase_06_idle_hibernate/
规划完成 Phase 6 软退出 RTC 机制。替代旧的"40s 硬重启退出"方案。

## 核心变更

### 1. 倒计时刷新（B+C 双源方案）

| 方案 | 监听源 | 实施位置 | 状态 |
|------|--------|---------|------|
| A 扬声器流 | I2S/PCM 输出 | application.cc audio output 3 处 | **宏关闭**（PHASE6_ENABLE_AUDIO_FALLBACK） |
| **B 字幕监听** | RTC subtitle 消息 | application.cc:1300 subtitle 分支 | **启用** |
| **C 智能体状态** | RTC conv_status 消息 | application.cc:1260 conv_status 分支 | **启用** |

复用现有 DIALOG_IDLE_COUNTDOWN_SECONDS=40 不新增常量。

### 2. 真退出 RTC 房间（释放 License）

- 新增 Protocol 基类虚函数 LeaveRoom（默认回退到 CloseAudioChannel）
- VolcRtcProtocol::LeaveRoom 覆写：volc_rtc_stop + volc_rtc_destroy
  - 火山官方文档明确：真退房必须 leaveRoom + destroyRTCEngine
  - CloseAudioChannel 只 stop 不够（真人仍在房间继续计费）
- 服务端 AI 任务在 180s 内自动清理（火山平台机制）

### 3. EnterIdleHibernate / WakeFromHibernate

EnterIdleHibernate 流程（严格顺序）：
1. protocol_->LeaveRoom()                  # 真退房
2. codec->EnableInput/Output(false)        # 重置 codec 状态机
3. recorder_pipeline_close()
4. hibernating_.store(true)                # 关键：先设标志阻止 PowerSaveTimer
5. esp_pm_configure(light_sleep=false)     # 双保险禁用 Light Sleep
6. SetDeviceState(kDeviceStateIdle)
7. idle_cycles_++ + NVS 持久化
8. 字幕"已自动退出RTC对话，按BOOT键重新连接RTC"（5 次重试间隔 200ms）

WakeFromHibernate 流程：
1. 检查 idle_cycles_ >= 50 → 硬重启清理碎片（兜底）
2. 清空字幕
3. ToggleChatState → OpenAudioChannel → 自动重建 rtc_handle_
4. RTC 重新加入房间（实测 2-3s 完成）

### 4. CanEnterSleepMode 加 hibernating 检查

防止 hibernate 期间 PowerSaveTimer 触发 esp_pm_configure(light_sleep=true)
导致 I2C 总线进入低功耗 → 唤醒后 ES7210/ES8311 通信失败 abort。

### 5. Dialog Watchdog 触发动作改造

旧：esp_restart() 整机重启（黑屏 15-25s + WiFi 重连）
新：Schedule(EnterIdleHibernate) 软退房（不熄屏 + 字幕提示）

### 6. BOOT 唤醒走 WakeFromHibernate 路径

iot_button 回调中检测 IsHibernating()，派发到独立 task 执行
WakeFromHibernate（避免阻塞 esp_timer 任务，CLAUDE.md 经验）。

### 7. OpenAudioChannel 适配重建

LeaveRoom 销毁 rtc_handle_ 后，OpenAudioChannel 头部检测 NULL
触发 Start() 异步重建，轮询 5s 等待就绪。NVS 缓存 device_secret
所以重建通常 100ms 完成。

## 实测验证（用户协作）

| 阶段 | 时间 |
|------|------|
| 40s 触发软休眠 | ✅ |
| LeaveRoom 真退房 | ✅ "✓ 已真退出 RTC 房间（leaveRoom + destroyRTCEngine）" |
| 屏幕保持 + 字幕显示 | ✅ "已自动退出RTC对话，按BOOT键重新连接RTC" |
| BOOT 按键唤醒 | ✅ |
| RTC 实例重建 | ✅ 100ms |
| RTC 重新加入房间 | ✅ 2-3s |
| 连续 2 次软休眠+唤醒 | ✅ 无 abort/I2C 失败 |
| 时间对比 | 旧硬重启 15-25s → 软休眠 3-5s（省 80%） |

## 6 个关键踩坑修复（详见 HIBERNATE_REPORT.md）

1. codec 状态机未重置 → 唤醒后 I2C abort
2. PowerSaveTimer Light Sleep 干扰 I2C 总线
3. hibernating_ 设置时序错误
4. dynamic_cast 在 -fno-rtti 下编译失败 → 改基类虚函数
5. LeaveRoom 后 OpenAudioChannel 直接失败 → 加重建逻辑
6. 字幕 LVGL 锁竞争 → 推迟到最后 + 5 次重试

## 文档产出（同时提交）

- .planning/.../phase_06_idle_hibernate/PLAN.md（含实施变更记录 V1-V6）
- .planning/.../phase_06_idle_hibernate/HIBERNATE_REPORT.md（验证报告）
- .planning/.../ROADMAP.md（Phase 1-5 ✅ + Phase 6 进行中状态更新）
- docs/Rtc_AIavatar/数字人表情渲染方案_云端预渲染+BLE+OTA.md
  新增第 19 章 RTC 空闲倒计时方案选型与软退出（9 小节）
- docs/Rtc_AIavatar/RTC软退出方案_移植参考.md
  完整移植参考（10 章 + 3 附录，可移植到其他火山 RTC 项目）
- docs/Rtc_AIavatar/音频卡顿_全局资源分析.md
  全局资源分析 + 13 项优化建议（不改代码）

2026-05-13 17:28:36 +08:00

.cache/clangd/index

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

.planning/milestones/digital_human_rtc

feat(rtc-only): Phase 6 - RTC 空闲软休眠（B+C 双源 + 真退房 + 字幕提示 + 内存兜底）

2026-05-13 17:28:36 +08:00

.vscode

feat(Rtc_AIavatar): 数字人透明 GIF 显示方案 PoC 完成（背景图+透明GIF叠加）

2026-05-12 17:14:49 +08:00

audios_new_p3

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

audios_p3

初始化补充提交

2026-02-24 15:57:32 +08:00

components

初始化补充提交

2026-02-24 15:57:32 +08:00

docs

feat(rtc-only): Phase 6 - RTC 空闲软休眠（B+C 双源 + 真退房 + 字幕提示 + 内存兜底）

2026-05-13 17:28:36 +08:00

dzbj @ 9223fd5a7d

补充提交dzbj文件夹

2026-02-27 10:44:58 +08:00

esp-spot

初始化补充提交

2026-02-24 15:57:32 +08:00

main

feat(rtc-only): Phase 6 - RTC 空闲软休眠（B+C 双源 + 真退房 + 字幕提示 + 内存兜底）

2026-05-13 17:28:36 +08:00

managed_components

feat: GIF动画表情系统 + 情绪映射增强 + HTTPS音频中止修复

2026-03-19 15:28:14 +08:00

scripts

初始化补充提交

2026-02-24 15:57:32 +08:00

spiffs_image

feat(rtc-only): Phase 3 - 数字人 GIF 资源准备（hiyori m03/m06/m07，209x360）

2026-05-13 11:42:30 +08:00

tests

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

tools

feat(rtc-only): Phase 3 - 数字人 GIF 资源准备（hiyori m03/m06/m07，209x360）

2026-05-13 11:42:30 +08:00

.gitignore

feat: 应援灯防撕裂优化 - DMA直接填充GRAM + LVGL flush拦截 + PWM黑屏遮蔽

2026-03-30 15:18:41 +08:00

00Kapi_Rtc_火山RTC整合移植方案.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

01Kapi_Rtc_WebSocket_替换为_火山RTC_技术分析报告.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

02Kapi_Rtc_火山RTC替换实现方案.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

03AEC_VOICE_INTERRUPT_PORTING_PLAN.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

04-2025-11-21音频优化记录.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

05-最新日志.txt

fix: 滑动切换图片时自动跳过解码失败的无效图片

2026-03-24 18:22:53 +08:00

06-AI对话和电子吧唧双模式适配说明 copy.md

docs: 更新 Claude Code 插件指南 + 新增开发参考文档

2026-04-02 10:09:08 +08:00

AEC_VAD_OPTIMIZATION.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

audios_new_p3.zip

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

BLE_JSON_通讯模块开发计划.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

BLE图片传输问题分析与优化建议.md

feat: 启用 BLE 5.0 2M PHY 图传加速 + 移除未使用的 BluFi 组件 + BLE 断连内存泄漏修复

2026-03-24 17:12:35 +08:00

BluFi蓝牙配网小程序开发需求说明书.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

BOOT_BUTTON_IMPLEMENTATION_COMPARISON.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

BOOT_BUTTON_LISTENING_STATE_IMPLEMENTATION_TEST.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

BOOT_BUTTON_MODIFICATION_SUMMARY.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

BOOT_BUTTON_NEW_IMPLEMENTATION_TEST.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

build.log

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

Claude Code 十个最值得装的 Skills copy.md

docs: 更新 Claude Code 插件指南 + 新增开发参考文档

2026-04-02 10:09:08 +08:00

Claude Code插件高效运用指南.md

chore: 更新插件指南/sdkconfig/IDF串口配置 + 从仓库移除 .DS_Store

2026-05-11 14:10:36 +08:00

CMakeLists.txt

feat: 完成 AI/吧唧双模式完全隔离重构 + 触摸坐标日志 + SPIFFS 预烧录

2026-02-28 10:23:04 +08:00

dependencies.lock

feat: 完成 AI/吧唧双模式完全隔离重构 + 触摸坐标日志 + SPIFFS 预烧录

2026-02-28 10:23:04 +08:00

idf_component.yml

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

ip_query_test.py

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

LICENSE

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

partitions_4M.csv

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

partitions_8M.csv

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

partitions_32M_sensecap.csv

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

partitions.csv

feat(rtc-only): Phase 2 - 16MB Flash 分区调整（OTA 5.5MB + SPIFFS 4.9MB）

2026-05-13 11:00:35 +08:00

postman请求.md

1、更新了postman参数、约束和语音请求讲故事的Function Call参数；

2026-03-02 17:27:03 +08:00

QMI8658A_IMU_Sensor_Development_Guide.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

QMI8658A驱动适配方案_B站驱动.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

QMI8658替换方案_Github驱动.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

README_en.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

README_ja.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

README_RTC.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

README.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

sdkconfig

feat(rtc-only): Phase 1 - 通过 CONFIG_BAJI_BADGE_MODE 屏蔽电子吧唧模式

2026-05-13 10:22:48 +08:00

sdkconfig.bak

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

sdkconfig.custom_wake_word

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

sdkconfig.defaults

feat(rtc-only): Phase 1 - 通过 CONFIG_BAJI_BADGE_MODE 屏蔽电子吧唧模式

2026-05-13 10:22:48 +08:00

sdkconfig.defaults.esp32c3

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

sdkconfig.defaults.esp32s3

1、启用LVGL GIF解码器（CONFIG_LV_USE_GIF=y），支持吧唧模式GIF图片BLE传输和播放；

2026-03-10 17:36:18 +08:00

sdkconfig.defaults.prod生产环境

feat: 启用 BLE 5.0 2M PHY 图传加速 + 移除未使用的 BluFi 组件 + BLE 断连内存泄漏修复

2026-03-24 17:12:35 +08:00

URGENT_INTERRUPT_FIX.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

VOICE_INTERRUPT_FEATURE.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

VOICE_INTERRUPT_OPTIMIZATION_GUIDE.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

volc_device_manager.o

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

和风天气运行日志.txt

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

自定义唤醒词移植说明.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

自定义唤醒词配置使用手册.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

蓝牙配网功能实现总结.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

蓝牙配网集成指南.md

1、初始化代码，待适配中....

2026-02-24 15:28:34 +08:00

资深嵌入式工程师开发思维.md

docs: 更新 Claude Code 插件指南 + 新增开发参考文档

2026-04-02 10:09:08 +08:00

README_en.md

XiaoZhi AI Chatbot

(中文 | English | 日本語)

Introduction

👉 Build your AI chat companion with ESP32+SenseVoice+Qwen72B!【bilibili】

👉 Equipping XiaoZhi with DeepSeek's smart brain【bilibili】

👉 Build your own AI companion, a beginner's guide【bilibili】

Project Purpose

This is an open-source project released under the MIT license, allowing anyone to use it freely, including for commercial purposes.

Through this project, we aim to help more people get started with AI hardware development and understand how to implement rapidly evolving large language models in actual hardware devices. Whether you're a student interested in AI or a developer exploring new technologies, this project offers valuable learning experiences.

Everyone is welcome to participate in the project's development and improvement. If you have any ideas or suggestions, please feel free to raise an Issue or join the chat group.

Learning & Discussion QQ Group: 376893254

Implemented Features

Wi-Fi / ML307 Cat.1 4G
BOOT button wake-up and interruption, supporting both click and long-press triggers
Offline voice wake-up ESP-SR
Streaming voice dialogue (WebSocket or UDP protocol)
Support for 5 languages: Mandarin, Cantonese, English, Japanese, Korean SenseVoice
Voice print recognition to identify who's calling AI's name 3D Speaker
Large model TTS (Volcano Engine or CosyVoice)
Large Language Models (Qwen, DeepSeek, Doubao)
Configurable prompts and voice tones (custom characters)
Short-term memory, self-summarizing after each conversation round
OLED / LCD display showing signal strength or conversation content
Support for LCD image expressions
Multi-language support (Chinese, English)

Hardware Section

Breadboard DIY Practice

See the Feishu document tutorial:

👉 XiaoZhi AI Chatbot Encyclopedia

Breadboard demonstration:

Supported Open Source Hardware

Firmware Section

Flashing Without Development Environment

For beginners, it's recommended to first use the firmware that can be flashed without setting up a development environment.

The firmware connects to the official xiaozhi.me server by default. Currently, personal users can register an account to use the Qwen real-time model for free.

👉 Flash Firmware Guide (No IDF Environment)

Development Environment

Cursor or VSCode
Install ESP-IDF plugin, select SDK version 5.3 or above
Linux is preferred over Windows for faster compilation and fewer driver issues
Use Google C++ code style, ensure compliance when submitting code

Developer Documentation

Board Customization Guide - Learn how to create custom board adaptations for XiaoZhi
IoT Control Module - Understand how to control IoT devices through AI voice commands

AI Agent Configuration

If you already have a XiaoZhi AI chatbot device, you can configure it through the xiaozhi.me console.

👉 Backend Operation Tutorial (Old Interface)

Technical Principles and Private Deployment

👉 Detailed WebSocket Communication Protocol Documentation

For server deployment on personal computers, refer to another MIT-licensed project xiaozhi-esp32-server

Star History

Languages

C 93.5%

C++ 1.9%

Jupyter Notebook 1.8%

Python 1.6%

Assembly 0.7%

Other 0.1%