关键词:chat_template, chat_chat_template, chat template
- 参考链接:
整体说明
- Chat Template 是为了适配不同模型产生的
- 理论上 Chat Template 与模型一一绑定,但同系列的模型经常相同
- Chat Template 定义一般在
tokenizer_config.json文件中的chat_template字段 - 使用 Chat Template 的函数是
AutoTokenizer.apply_chat_template,这是Hugging Face Transformers 库中 tokenizer 的一个核心方法 - 论文主要详细介绍
apply_chat_template函数的用法
apply_chat_template 函数签名
- 函数签名详情:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18def apply_chat_template(
self,
conversation: Union[list[dict[str, str]], list[list[dict[str, str]]]],
tools: Optional[list[Union[dict, Callable]]] = None,
documents: Optional[list[dict[str, str]]] = None,
chat_template: Optional[str] = None,
add_generation_prompt: bool = False,
continue_final_message: bool = False,
tokenize: bool = True,
padding: Union[bool, str, PaddingStrategy] = False,
truncation: bool = False,
max_length: Optional[int] = None,
return_tensors: Optional[Union[str, TensorType]] = None,
return_dict: bool = False,
return_assistant_tokens_mask: bool = False,
tokenizer_kwargs: Optional[dict[str, Any]] = None,
**kwargs,
) -> Union[str, list[int], list[str], list[list[int]], BatchEncoding]:
apply_chat_template 核心参数详解
conversation(必需参数)
- 对话数据,可以是单个对话列表或批量对话列表,每个消息必须包含
"role"和"content"键 - 参数类型:
Union[list[dict[str, str]], list[list[dict[str, str]]]] - 简单参考示例:
1
2
3
4
5
6
7
8
9
10
11
12
13# 单轮对话
conversation = [
{"role": "user", "content": "你好,请问今天天气如何?"},
{"role": "assistant", "content": "您好!我需要知道您的位置才能提供准确的天气信息。"}
]
# 多轮对话
conversation = [
{"role": "system", "content": "你是一个有用的AI助手"},
{"role": "user", "content": "解释量子计算"},
{"role": "assistant", "content": "量子计算是利用量子力学现象进行计算的技术..."},
{"role": "user", "content": "它与传统计算有什么区别?"}
]
tokenize
- 控制输出格式是否为 tokenized(token ID 列表)还是文本字符串
- 参数类型:
bool,默认值是True- 对于需要直接传递给模型的场景,建议保持
True; - 对于需要查看格式化文本或进行自定义处理的场景,可以设置为
False
- 对于需要直接传递给模型的场景,建议保持
- 简单参考示例:
1
2
3
4
5
6
7
8
9
10
11
12# 输出文本格式
text_output = tokenizer.apply_chat_template(
conversation,
tokenize=False
)
print(text_output) # 例如: <|user|>你好!<|end|><|assistant|>
# 输出token IDs
token_output = tokenizer.apply_chat_template(
conversation,
tokenize=True
)
add_generation_prompt
- 是否在格式化输出末尾添加助手回复的提示符
- 参数类型:
bool,**默认值:False - 这是确保模型正确生成助手回复的关键参数
- 当设置为
True时,会在对话末尾添加表示助手开始回复的特殊标记
- 当设置为
- 简单参考示例:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15# 不添加生成提示
output_without = tokenizer.apply_chat_template(
conversation,
add_generation_prompt=False,
tokenize=False
)
# 输出: <|user|>你好!<|end|>
# 添加生成提示
output_with = tokenizer.apply_chat_template(
conversation,
add_generation_prompt=True,
tokenize=False
)
# 输出: <|user|>你好!<|end|><|assistant|>
continue_final_message
- 是否继续最后一个消息而不是开始新消息
- 参数类型:
bool,默认值是False - 用于”预填充”模型回复的场景,当你想让模型继续已有的助手消息时使用
- 不能与
add_generation_prompt同时使用
- 不能与
- 简单参考示例:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16# 预填充助手回复
conversation = [
{"role": "user", "content": "请用JSON格式回答"},
{"role": "assistant", "content": '{"answer": "'}
]
# 模型将继续这个JSON格式,而不是开始新消息
formatted = tokenizer.apply_chat_template(
conversation,
continue_final_message=True,
tokenize=True
)
## continue_final_message=True,此时不拼接结束符号
# [Round 0] USER:请用JSON格式回答 ASSISTANT:{"answer": "
## continue_final_message=False(默认),此时拼接结束符号,表示本次会话已经结束
# [Round 0] USER:请用JSON格式回答 ASSISTANT:{"answer": "</longcat_s>
chat_template
- 自定义的 Jinja2 模板字符串
- 参数类型:
Optional[str],默认值为None - 当需要使用非默认模板或测试新模板格式时使用
- 简单参考示例:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15custom_template = """
{% for message in messages %}
{% if message['role'] == 'user' %}
Human: {{ message['content'] }}
{% elif message['role'] == 'assistant' %}
Assistant: {{ message['content'] }}
{% endif %}
{% endfor %}
"""
output = tokenizer.apply_chat_template(
conversation,
chat_template=custom_template,
tokenize=False
)
tools
工具列表,用于函数调用场景
参数类型:
Optional[list[Union[dict, Callable]]],默认值为None每个工具应该是 JSON Schema 格式,包含名称、描述和参数类型
简单参考示例:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "获取天气信息",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "城市名称"}
},
"required": ["location"]
}
}
}
]
formatted = tokenizer.apply_chat_template(
conversation,
tools=tools,
add_generation_prompt=True
)- 以上函数的调用方式如
get_weather(**call_info["arguments"])- 其中
call_info["arguments"]是一个必须包含"location"的字典
- 其中
- 以上函数的调用方式如
documents(待尝试)
- 文档列表,用于 RAG(检索增强生成)场景
- 参数类型:
Optional[list[dict[str, str]]],默认值为None- 推荐格式为 每个文档应包含
"title"和"text"键
- 推荐格式为 每个文档应包含
- 简单参考示例:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16documents = [
{
"title": "量子计算简介",
"text": "量子计算是利用量子力学现象进行计算的技术..."
},
{
"title": "量子比特",
"text": "量子比特是量子计算的基本单位..."
}
]
formatted = tokenizer.apply_chat_template(
conversation,
documents=documents,
add_generation_prompt=True
)
多模态 chat-template 使用详解
- 注:多模态中,Processor 组合了 Tokenizer + 图像处理
- 理论上,Processor 中已经包含了处理文本的 Tokenizer 了
简单示例:多模态下的模型生成示例(可本地执行)
模型生成示例(MacOS 也可执行)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62# 阻止 fla 被 transformers 动态导入(For MacOS,平时不需要这部分)
import sys
sys.modules["fla"] = None
try:
import fla
except ImportError:
pass
from transformers import AutoModelForCausalLM, AutoProcessor
from qwen_vl_utils import process_vision_info
from PIL import Image
# 1. 加载模型和 Processor
model_path = "/Users/sanye/llm/model/Qwen3.5-0.8B"
processor = AutoProcessor.from_pretrained(model_path)
# 加载到 CPU 中执行 (For MacOS)
model = AutoModelForCausalLM.from_pretrained(model_path, device_map="cpu", torch_dtype="float32")
# 2. 准备对话消息
messages = [
{
"role": "user",
"content": [
# 传入本地 image 对象
{"type": "image", "image": "/Users/sanye/llm/image/candy.JPG"},
{"type": "text", "text": "What animal is on the candy?"}
]
},
]
# 3. 预处理视觉信息 (关键步骤)
# 这一步会读取并处理图片,返回一个 PIL.Image 对象的列表
image_inputs, video_inputs = process_vision_info(messages)
# 4. 处理文本和视觉信息
# 首先,使用 apply_chat_template 生成纯文本部分
text = processor.apply_chat_template(
messages,
tokenize=False, # 先不 tokenize,只生成文本字符串
add_generation_prompt=True,
)
# 然后,将文本和图像一起传给 processor 进行最终的 tokenization 和处理
inputs = processor(
text=[text],
images=image_inputs,
videos=video_inputs, # 如果没有视频,传入 None 即可
padding=True,
return_tensors="pt",
).to(model.device)
print(inputs)
# 5. 生成输出
# 至此,inputs 可以直接作为 generate 的参数使用
outputs = model.generate(**inputs, max_new_tokens=40)
# 解码并打印生成的文本
print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
# # 输出类似下面的形式:
# Based on the image, there are **two animals** painted on the candies:
#
# 1. **A turtle** — painted on the teal candy in the upper part of the hand.qwen_vl_utils.process_vision_info的详细说明见本文附录- 所有函数的更多详细说明见下文
构造简单示例和详细解读
先生成可读 chat_template 模版,再进行 tokenize 的示例(简单图片示例)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112from transformers import AutoProcessor
model_path = "/Users/sanye/llm/model/Qwen3.5-0.8B"
processor = AutoProcessor.from_pretrained(model_path)
# 读取本地图片
from PIL import Image
# image = Image.open("/Users/sanye/llm/image/candy.JPG").convert("RGB")
import numpy as np
# 随机初始化一张 (256 x 256 x 3) 的图片,256 = 16 * 16,16 是 Qwen3.5 patch 像素切割单位(config.json 中的 patch_size), 16 是默认最小 patch 数(小于 16 也会填充到 16)
image = Image.fromarray(np.random.randint(0, 255, (256, 256, 3), dtype=np.uint8))
messages = [
{
"role": "user",
"content": [
{"type": "text", "text": "We have two images"},
{"type": "text", "text": "\n# image1:"},
{"type": "image", "image": image},
{"type": "text", "text": "\n# image2:"},
{"type": "image", "image": image},
{"type": "text", "text": "\nWhat animal is on the candy?"},
]
},
]
prompt = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=False
)
# 输出结果见末尾
print(prompt)
from qwen_vl_utils import process_vision_info
# 预处理视觉信息,这一步会下载并处理图片和视频(若存在),返回 PIL.Image 对象的列表
images, videos = process_vision_info(messages)
# 等价于:images = [image, image]
inputs = processor(
text=prompt,
images=images,
return_tensors="pt"
)
# 下面得到的结果和前面的示例一模一样
print(inputs)
for key in inputs:
value = inputs[key]
print(f"key: {key}")
print(f"value: {value.shape}")
# <|im_start|>user
# We have two images
# # image1:<|vision_start|><|image_pad|><|vision_end|>
# # image2:<|vision_start|><|image_pad|><|vision_end|>
# What animal is on the candy?<|im_end|>
# <|im_start|>assistant
# <think>
#
# </think>
#
#
# {'input_ids': tensor([[248045, 846, 198, 1596, 599, 1330, 5167, 198, 2,
# 2099, 16, 25, 248053, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248054, 198, 2, 2099,
# 17, 25, 248053, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056, 248056,
# 248056, 248056, 248056, 248056, 248054, 198, 3710, 9572, 369,
# 383, 279, 30517, 30, 248046, 198, 248045, 74455, 198,
# 248068, 271, 248069, 271]]),
# 'attention_mask': tensor([[1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
# 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
# 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
# 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
# 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
# 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
# 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1]]),
# 'mm_token_type_ids': tensor([[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
# 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
# 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
# 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
# 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
# 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
# 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]]),
# 'pixel_values': tensor([[ 0.9059, 0.8980, 0.8745, ..., -0.7725, -0.7725, -0.7333],
# [ 0.0431, 0.1765, 0.3176, ..., 0.4353, 0.4667, 0.4824],
# [ 0.0745, 0.0667, 0.0588, ..., 0.4667, 0.3961, 0.3098],
# ...,
# [ 0.1451, 0.1373, 0.1137, ..., -0.7412, -0.7569, -0.7647],
# [-0.4431, -0.4902, -0.5294, ..., 0.2000, -0.0745, -0.3255],
# [-0.6157, -0.5843, -0.5686, ..., 0.5373, 0.5686, 0.5843]]),
# 'image_grid_thw': tensor([[ 1, 16, 16],
# [ 1, 16, 16]])}
#
# key: input_ids
# value: torch.Size([1, 166])
# key: attention_mask
# value: torch.Size([1, 166])
# key: mm_token_type_ids
# value: torch.Size([1, 166])
# key: pixel_values
# value: torch.Size([512, 1536])
# key: image_grid_thw
# value: torch.Size([2, 3])processor.apply_chat_template说明:- 这里可视化的结果中,将每一张图片都作为一个
placeholder( Qwen3.5 中是<|vision_start|><|image_pad|><|vision_end|>)专门存放起来了- 整个输出的
prompt是可读的
- 整个输出的
- 后续在进行 tokenize 时,需要传入这个 Prompt 和 image 信息(对齐顺序),从而保证 tokenize 后将图片 Token 正确插入到指定的顺序中
- 这里可视化的结果中,将每一张图片都作为一个
说明:
inputs对象包含了处理多模态输入(文本 + 图像)后得到的张量,各字段含义如下:input_ids:- Shape:
(1, seq_len)- 示例中
seq_len = 166
- 示例中
- 经过 tokenizer 编码后的完整提示文本(包含特殊标记如
<|vision_start|>(248053)、<|image_pad|>(248056) 、<|vision_end|>(248054)等)的 token ID 序列 - 这里的长度和 Prompt 是完全对不齐的,多出来的 Token 就是图片转换成的 Token 占位符
- 注意:这里的图片 Token 占位符的 连续 64 个 Token 都是 248056 (
<|image_pad|>),本质上不包含任何数据 - 注意:这里的 Token 数和后面对应的 patch 数并不相等(详情见
image_grid_thw中 每张图片被 ViT 编码成 256 个 Token), 因为还会压缩合并 Token- 通过 空间合并(Spacial Merging)技术,patch 经过压缩后得到的数字才是结果
- 通过模型结构中的 PatchMerger 模块(一般是 MLP)来合并相邻的 多个 (n x n) Token 为一个 Token,一般都选择
spatial_merge_size=2(即长宽各合并 2 个 Token 合并(2x2)为一组)
- 通过模型结构中的 PatchMerger 模块(一般是 MLP)来合并相邻的 多个 (n x n) Token 为一个 Token,一般都选择
- 比如这里经过了
spatial_merge_size=2,最终得到 256/(2x2)=64 个 最终的 Token - 由于目前大部分模型(比如 Kimi K2.5 等)都不需要生成 图片或视频,所以不涉及到解码(NTP 训练时,视觉 Token 对应的位置不会有损失回传)
- 最后在模型中真正参与 Attention 的是这些 Merging 后的 Token,但是为了保证 Patch 位置可识别,会在 2D-NoPE 上添加补偿
- 通过 空间合并(Spacial Merging)技术,patch 经过压缩后得到的数字才是结果
- 注意:这里的图片 Token 占位符的 连续 64 个 Token 都是 248056 (
- Shape:
attention_mask:- Shape: 与
input_ids相同 - 标记每个位置是否为有效 token(
1)或填充(0),用于自注意力机制中忽略填充部分
- Shape: 与
mm_token_type_ids:- Shape: 与
input_ids相同 - 用于区分 token 的类型,例如文本 token 和视觉 token(图像占位符),方便模型在多模态处理时识别不同模态的输入
- Shape: 与
pixel_values(ViT 编码后的结果) :- Shape:
(total_patches, feature_dim)- 在本例中为
(512, 1536)
- 在本例中为
- 它并不是原始的像素矩阵,而是 所有图像经过视觉编码器(如 ViT)分割成 patch 并投影后的特征向量 ,按顺序拼接在一起
total_patches = 512等于所有图像的 patch 总数 ,feature_dim = 1536是每个 patch 的 Embedding 维度- 如何区分不同图片 :
pixel_values本身只是线性拼接,不直接携带图片归属信息- 需要结合
image_grid_thw来确定每张图片的 patch 数量- 理解:
pixel_values将多张图片的 patch 特征按传入顺序拼接为单张二维张量(单从这里无法切分多张图片的内容)image_grid_thw记录了每张图片的网格结构,通过计算每张图片的 patch 总数,即可从pixel_values中切分出每张图片的特征,从而区分不同图片
- 理解:
- Shape:
image_grid_thw(ViT 编码 Token 的信息):- Shape
(num_images, 3)- 本例为
(2, 3),表示有num_images=2张图片
- 本例为
- 每张图片的 3 维时空网格尺寸
(temporal, height, width),本示例中每张图片的值为[1, 16, 16]temporal是时间维度- 对于静态图像,
temporal通常为1 - 对于视频,是根据
temporal ≈ num_frames / temporal_patch_size来确定num_frames:模型实际采样并处理的视频帧数,模型不会处理视频的每一帧,而是会进行采样temporal_patch_size:时序上的“合并”尺寸- Qwen2.5-VL默认将连续的 2 帧合并为 1 个“时序补丁”(temporal patch)
- 即
temporal_patch_size=2,配置在 config.json 中
- 即
- 这是实现时序维度的 PatchMerger
- Qwen2.5-VL默认将连续的 2 帧合并为 1 个“时序补丁”(temporal patch)
- 对于静态图像,
height和width是图像被划分的 patch 网格数(例如本例中为16x16)- 注意:这里已经没有了传统图片的
Channel维度,在编码为 Patch 时,图片的 channel 维度信息已经被编码到1536维度的向量中了
- 每张图片的 patch 总数 =
temporal × height × width- 本例中每张图片有
1 × 16 × 16 = 256个 patch,因此pixel_values的前256行对应第一张图片,后256行对应第二张图片(有序索引)
- 本例中每张图片有
- Shape
补充:更多真实图片下的示例
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81from transformers import AutoProcessor
model_path = "/Users/sanye/llm/model/Qwen3.5-0.8B"
processor = AutoProcessor.from_pretrained(model_path)
# 读取本地图片
from PIL import Image
image = Image.open("/Users/sanye/llm/image/candy.JPG").convert("RGB")
messages = [
{
"role": "user",
"content": [
{"type": "text", "text": "We have two images"},
{"type": "text", "text": "\n# image1:"},
{"type": "image", "image": image},
{"type": "text", "text": "\n# image2:"},
{"type": "image", "image": image},
{"type": "text", "text": "\nWhat animal is on the candy?"},
]
},
]
prompt = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=False
)
# 输出结果见末尾
print(prompt)
from qwen_vl_utils import process_vision_info
# 预处理视觉信息,这一步会下载并处理图片和视频(若存在),返回 PIL.Image 对象的列表
images, videos = process_vision_info(messages)
inputs = processor(
text=prompt,
images=images,
return_tensors="pt"
)
# 下面得到的结果和前面的示例一模一样
print(inputs)
for key in inputs:
value = inputs[key]
print(f"key: {key}")
print(f"value: {value.shape}")
# # print(prompt) 输出如下:
# <|im_start|>user
# We have two images
# # image1:<|vision_start|><|image_pad|><|vision_end|>
# # image2:<|vision_start|><|image_pad|><|vision_end|>
# What animal is on the candy?<|im_end|>
# <|im_start|>assistant
# <think>
#
# </think>
#
# # print(inputs) 输出如下:
# {'input_ids': tensor([[248045, 846, 198, ..., 271, 248069, 271]]),
# 'attention_mask': tensor([[1, 1, 1, ..., 1, 1, 1]]),
# 'mm_token_type_ids': tensor([[0, 0, 0, ..., 0, 0, 0]]),
# 'pixel_values': tensor([[ 0.4588, 0.5059, 0.5294, ..., 0.5373, 0.5608, 0.5686],
# [ 0.5373, 0.5137, 0.5137, ..., 0.5216, 0.5451, 0.5216],
# [ 0.5451, 0.5451, 0.5451, ..., 0.5922, 0.5686, 0.5608],
# ...,
# [-0.3020, -0.2549, -0.2706, ..., -0.1765, -0.1686, -0.2000],
# [-0.1922, -0.1451, -0.1216, ..., -0.0275, -0.1059, -0.0980],
# [-0.0510, -0.0510, -0.0588, ..., -0.1059, -0.1059, -0.1216]]),
# 'image_grid_thw': tensor([[ 1, 188, 252],
# [ 1, 188, 252]])}
#
# key: input_ids
# value: torch.Size([1, 23726])
# key: attention_mask
# value: torch.Size([1, 23726])
# key: mm_token_type_ids
# value: torch.Size([1, 23726])
# key: pixel_values
# value: torch.Size([94752, 1536])
# key: image_grid_thw
# value: torch.Size([2, 3])- 解读:
pixel_values:- Shape:
(total_patches, feature_dim)- 在本例中为
(94752, 1536)
- 在本例中为
- 它并不是原始的像素矩阵,而是所有图像经过视觉编码器(如 ViT)分割成 patch 并投影后的特征向量 ,按顺序拼接在一起
total_patches = 94752等于所有图像的 patch 总数 ,feature_dim = 1536是每个 patch 的 Embedding 维度- 如何区分不同图片 :
pixel_values本身只是线性拼接,不直接携带图片归属信息- 需要结合
image_grid_thw来确定每张图片的 patch 数量
- Shape:
image_grid_thw:- Shape
(num_images, 3)- 本例为
(2, 3),表示有num_images=2张图片
- 本例为
- 每张图片的 3 维时空网格尺寸
(temporal, height, width),本示例中每张图片的值为[1, 188, 252]- 对于静态图像,
temporal通常为1 height和width是图像被划分的 patch 网格数(例如本例中为188×252)
- 对于静态图像,
- 每张图片的 patch 总数 =
temporal × height × width- 本例中每张图片有
1 × 188 × 252 = 47376个 patch,因此pixel_values的前47376行对应第一张图片,后47376行对应第二张图片(有序索引)
- 本例中每张图片有
- Shape
- 解读:
更多合并的编码示例(注:不建议合并使用)
- 注:建议先 apply_chat_template 为可读的文本,然后再进行 Tokenize 更合适
- 合并编码(仅调用一次 processor,看似更简单,但是不推荐)情况展示
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49from transformers import AutoProcessor
model_path = "/Users/sanye/llm/model/Qwen3.5-0.8B"
processor = AutoProcessor.from_pretrained(model_path)
messages = [
{
"role": "user",
"content": [
{"type": "image", "image": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
{"type": "image", "image": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
{"type": "text", "text": "What animal is on the candy?"}
]
},
]
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
)
print(inputs)
for key in inputs:
value = inputs[key]
print(f"key: {key}")
print(f"value: {value.shape}")
# {'input_ids': tensor([[248045, 846, 198, ..., 271, 248069, 271]]),
# 'attention_mask': tensor([[1, 1, 1, ..., 1, 1, 1]]),
# 'mm_token_type_ids': tensor([[0, 0, 0, ..., 0, 0, 0]]),
# 'pixel_values': tensor([[ 0.4588, 0.5059, 0.5294, ..., 0.5373, 0.5608, 0.5686],
# [ 0.5373, 0.5137, 0.5137, ..., 0.5216, 0.5451, 0.5216],
# [ 0.5451, 0.5451, 0.5451, ..., 0.5922, 0.5686, 0.5608],
# ...,
# [-0.3020, -0.2549, -0.2706, ..., -0.1765, -0.1686, -0.2000],
# [-0.1922, -0.1451, -0.1216, ..., -0.0275, -0.1059, -0.0980],
# [-0.0510, -0.0510, -0.0588, ..., -0.1059, -0.1059, -0.1216]]),
# 'image_grid_thw': tensor([[ 1, 188, 252],
# [ 1, 188, 252]])}
#
# key: input_ids
# value: torch.Size([1, 23711])
# key: attention_mask
# value: torch.Size([1, 23711])
# key: mm_token_type_ids
# value: torch.Size([1, 23711])
# key: pixel_values
# value: torch.Size([94752, 1536])
# key: image_grid_thw
# value: torch.Size([2, 3])
输出控制参数
return_tensors
- 指定返回的张量类型,
return_tensors参数仅在tokenize=True时生效 - 为了同时支持多种框架,特意设计了不同的返回值类型,减少一次使用时重新转类型的时间,提升速度
- 参数类型:
Optional[Union[str, TensorType]]- 默认值为
None,此时返回 原生 Python 数据结构 - 可选值:
'pt'(PyTorch),'tf'(TensorFlow),'np'(NumPy),'jax'(JAX)
- 默认值为
- 返回张量类型详细说明:
- 若
tokenize=False,则无论如何返回值都是原生 Python 数据结构 - 若
tokenize=True,则根据return_tensors判断返回类型- 默认值为
None,此时返回 原生 Python 数据结构 - 指定时根据上述指定类型返回值
- 默认值为
- 若
- 简单参考示例:
1
2
3
4
5
6
7
8
9
10
11
12
13# PyTorch tensors
pytorch_output = tokenizer.apply_chat_template(
conversation,
return_tensors="pt",
add_generation_prompt=True
)
# NumPy arrays
numpy_output = tokenizer.apply_chat_template(
conversation,
return_tensors="np",
add_generation_prompt=True
)
return_dict
- 是否返回包含多个字段的字典,仅在
tokenize=True时有效 - 参数类型:
bool,默认值为False - 当
apply_chat_template函数的return_dict参数设为True时,函数会返回一个字典格式的结果 ,而非默认的张量(或字符串)- 默认(
return_dict=False):若tokenize=True,直接返回input_ids张量(若开启 padding,会返回形状为[batch_size, seq_len]的张量) return_dict=True:返回字典,键值对应模型输入的核心要素,结构清晰,无需手动区分张量类型
- 默认(
- 字典中会根据需要明确包含
input_ids、attention_mask等模型输入所需的关键组件,便于直接拆解和使用- 当
return_assistant_tokens_mask=True时,会多返回一个assistant_masks字段
- 当
- 具体来说,返回的
output是一个字典,包含以下关键键(根据参数配置可能增减):"input_ids":对话文本对应的token ID张量(模型输入核心)"attention_mask":注意力掩码张量(标记哪些token是有效文本,哪些是padding,避免模型关注padding)
return_assistant_tokens_mask
参数类型:
bool,默认值为False当
return_assistant_tokens_mask=True时,会多返回一个assistant_masks字段是否返回助手生成 token 的掩码,助手生成的 token 对应掩码为 1,用户和系统 token 对应掩码为 0
- 只在
tokenize=True且return_dict=True时生效,否则返回错误:1
ValueError: `return_dict=True` is incompatible with `tokenize=False`, because there is no dict of tokenizer outputs to return.
- 只在
仅支持包含
{% generation %} ... {% endgeneration %}关键字的聊天模板- 这个关键字不会影响正常的渲染,只是用于标记
assistant的内容 - 一般来说,用该标记将
assistant的整个内容(包括应该学习的所有内容,如 tools 调用等内容都包括进来) - 注意:原始的 Jinja 语法是不可以随便写这种位置标签的,但是 apply_chat_template 会对 这个标签做特殊处理,所以不用担心
- 验证:若随机增加未知的标签
{% generationa %} ... {% endgenerationa %}则会出现下面的问题:1
jinja2.exceptions.TemplateSyntaxError: Encountered unknown tag 'generationa'. Jinja was looking for the following tags: 'elif' or 'else' or 'endif'. The innermost block that needs to be closed is 'if'.
- 验证:若随机增加未知的标签
- 这个关键字不会影响正常的渲染,只是用于标记
亲测:
return_assistant_tokens_mask=True但对 chat_template 有一定的要求(要求包含{% generation %}用于标记assistant位置),当前大部分 chat_template 都不支持,此时会全屏蔽(返回assistant_masks字段全是 0)- 警告信息如下:
1
return_assistant_tokens_mask==True but chat template does not contain `{% raw %} {% generation %} {% endraw %}` keyword.
- 警告信息如下:
常用用途:用于训练或分析,标识哪些 token 是由助手生成的
简单参考示例:
1
2
3
4
5
6
7
8output = tokenizer.apply_chat_template(
conversation,
return_dict=True,
tokenize=True,
return_assistant_tokens_mask=True,
add_generation_prompt=True
)
# 输出包含助手token掩码信息- 返回结果
output多多包涵一个assistant_masks的list类型字段,里面为 1 的地方都是assistant的Token - 需要注意返回结果中为 1 的 Token 中,可能会多出来一些自定义的 Token,此时需要手动处理一下
- 比如在
assistant信息前面加入的<USER>Token 通常会被包含 - 此时需要自己手动识别并去除一下相关的特殊 Token
- 比如在
- 返回结果
序列处理参数
padding
- 用于指定填充类型
- 参数类型:
Union[bool, str, PaddingStrategy],默认值为False - 仅在
tokenize=True时生效- 虽然说明文档中未明确说明这一点,但经过测试,
tokenize=False时不会 padding
- 虽然说明文档中未明确说明这一点,但经过测试,
- 可选值为
True或'longest': 填充到批次中最长序列'max_length': 填充到指定最大长度,最大长度由 max_length 参数指定False(默认值) 或'do_not_pad': 不填充
- 简单参考示例:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16# 填充到最长序列
padded_output = tokenizer.apply_chat_template(
batch_conversations,
padding=True,
tokenize=True,
return_tensors="pt"
)
# 填充到最大长度
max_length_output = tokenizer.apply_chat_template(
conversation,
padding="max_length",
max_length=512,
tokenize=True,
return_tensors="pt"
)
补充:padding_side 属性指定左填充 or 右填充
Tokenizer.apply_chat_template方法本身并不直接提供选择左 padding(Left Padding) 或右 padding (Right Padding) 的参数- padding 方式通常是由 tokenizer 的整体配置(特别是
padding_side参数)决定的,而不是由apply_chat_template方法单独控制- 如果需要设置 padding 方向,应该在初始化 tokenizer 时或通过
tokenizer.padding_side属性进行配置 padding_side可选值为"right"(默认值) 或"left"
- 如果需要设置 padding 方向,应该在初始化 tokenizer 时或通过
- 注:
apply_chat_template方法主要用于将对话历史格式化为模型期望的输入格式,它会调用 tokenizer 的编码逻辑,而编码过程会遵循 tokenizer 已设置的padding_side配置 - 示例代码:
1
2
3
4
5
6
7
8
9
10
11
12
13
14from transformers import AutoTokenizer
# 初始化 tokenizer 时指定 padding 方向
tokenizer = AutoTokenizer.from_pretrained("model_name", padding_side="left")
# 核心代码
tokenizer.padding_side = "right"
# 应用对话模板时会遵循上述 padding 配置
chat = [
{"role": "user", "content": "Hello!"},
{"role": "assistant", "content": "Hi there!"}
]
inputs = tokenizer.apply_chat_template(chat, tokenize=True, return_tensors="pt", padding=True)
truncation
- 是否截断超过最大长度的序列
- 在处理长对话时启用,避免超出模型最大长度限制
- 参数类型:
bool,默认值为False
补充:truncation_side 属性指定左截断 or 右截断
- 用法与
padding_side参数类似 truncation_side是Tokenizer类的一个属性,其可选值为:"left":从序列的左侧(开头)截断"right"(默认值):从序列的右侧(结尾)截断(默认值)
max_length
- 最大长度限制(以token数计),与
padding或truncation配合使用 - 参数类型:
Optional[int],默认值为None
tokenizer_kwargs
- 传递给分词器的额外参数,类型为 Optional[dict[str, Any]]
**kwargs
- 传递给模板渲染器的额外参数,可在聊天模板中访问
附录:完整使用示例
基础对话生成
- 简单对话简单示例:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
# 加载模型和tokenizer
model_id = "xxx/xxx"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
# 准备对话
messages = [
{"role": "system", "content": "你是一个友好的AI助手"},
{"role": "user", "content": "请解释机器学习的基本概念"},
]
# 格式化对话
tokenized_chat = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt"
)
# 生成回复
with torch.no_grad():
outputs = model.generate(
tokenized_chat,
max_new_tokens=256,
temperature=0.7,
do_sample=True
)
# 解码回复
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
批量处理
- 多个对话同时处理
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26# 批量对话
batch_conversations = [
[
{"role": "user", "content": "什么是人工智能?"}
],
[
{"role": "user", "content": "如何学习编程?"}
]
]
# 批量格式化
batch_output = tokenizer.apply_chat_template(
batch_conversations,
padding=True,
truncation=True,
max_length=512,
return_tensors="pt",
add_generation_prompt=True
)
# 批量生成
batch_outputs = model.generate(
**batch_output,
max_new_tokens=128,
temperature=0.7
)
RAG场景使用(待补充)
RAG 使用模板的示例
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24# 准备文档
documents = [
{
"title": "2025年科技趋势报告",
"text": "人工智能和机器学习技术将继续快速发展..."
},
{
"title": "量子计算进展",
"text": "量子计算机在特定问题上展现出巨大优势..."
}
]
# 用户问题
conversation = [
{"role": "user", "content": "2025年有哪些重要的科技趋势?"}
]
# 使用RAG模板格式化
formatted_input = tokenizer.apply_chat_template(
conversation,
documents=documents,
add_generation_prompt=True,
return_tensors="pt"
)注:一般的模板不支持 documents,目前包括 Llama 系列,Qwen 系列等均不支持
支持
documents的模型示例:huggingface.co/CohereLabs/c4ai-command-r-v01/blob/main/tokenizer_config.jsonchat_template模版不支持documents时,方法包括:- 自己写 Jinja2 模板,把
documents拼进system或第一条user消息,再在 vLLM 等框架启动时通过--chat-template指定,apply_chat_template函数也支持该参数; - 直接在外部把检索结果拼接成普通字符串,再按常规
messages=[{"role":"user","content":"..."}]传入即可
- 自己写 Jinja2 模板,把
最佳实践:先通过 Prompt Engineering 找到合适的模版,然后通过固定的模版文件将该形式固定下来,模型微调和线上 serving 均使用这个模版,这样可以避免因为模型微调和线上 serving 不一致带来的问题,也方便团队内外的合作
附录:最佳实践建议
add_generation_prompt的使用: 在推理时,确保设置add_generation_prompt=True以获得正确的助手回复- 不同模型使用不同的 chat template,可使用
tokenizer.get_chat_template()查看具体格式 - 对于长对话,启用
truncation=True并设置合适的max_length - 批量处理时合理使用
padding参数以提高效率,否则可能返回不同长度的编码结果 - 添加适当的错误处理,特别是对于模板不支持的功能
- 某些模型不支持
tools参数,需要检查模型文档 - 处理长序列时可能遇到内存问题,考虑减小
batch size或max_length - 确保
conversation格式正确,每个消息都有role和content键
附录:关于 tools 类型
- 在大模型工具调用场景中,
code_interpreter(代码解释器)和function(函数/工具)是两种不同类型的工具
function(函数/工具)
function(函数/工具)是预先定义的、具有特定功能的程序函数或API接口,用于让模型调用外部能力 ,模型通过生成符合格式的调用指令(如JSON),触发这些函数执行,并获取返回结果- 例如:天气查询接口、数据库查询函数、网页爬虫工具等,如模型调用
get_weather(city="北京")函数获取实时天气
- 例如:天气查询接口、数据库查询函数、网页爬虫工具等,如模型调用
function调用方式:模型需严格按照预设格式(如{"name": "函数名", "parameters": {"参数名": "值"}})生成调用指令,确保函数能被正确解析和执行,例如:1
{"name": "translate", "parameters": {"text": "Hello", "target_lang": "zh"}}
function灵活性低:功能固定,只能执行预定义的操作;但安全性高:严格限制在预设函数范围内,风险可控function适用场景包括- 需要调用外部服务或系统(如查询实时数据、操作硬件设备)
- 执行结构化任务(如数据库查询、API调用)
- 功能固定、无需动态逻辑的操作(如格式转换、简单计算)
code_interpreter(代码解释器)
code_interpreter(代码解释器)是一个能够动态执行代码(通常是Python)的沙箱环境,允许模型直接生成并运行代码来解决问题,模型生成代码后,解释器会运行代码并返回输出结果(包括文本、图表等)- 例如:执行数学计算、数据可视化、处理 Excel 表格等,如 模型生成Python代码计算
1+2+...+100的和,并通过代码解释器执行得到结果
- 例如:执行数学计算、数据可视化、处理 Excel 表格等,如 模型生成Python代码计算
code_interpreter灵活性高:支持任意代码逻辑,可解决复杂、动态的问题;但安全性低:需运行用户/模型生成的代码,存在恶意代码风险(通常通过沙箱隔离缓解)code_interpreter中,模型直接生成代码片段(通常包裹在特定标记中,如... ```),由解释器解析并运行,例如: 1
2
3
4```python
import numpy as np
result = np.sum(range(1, 101))
print(result)code_interpreter适用场景包括- 需要复杂逻辑计算(如统计分析、公式推导)
- 数据处理与可视化(如绘制图表、处理CSV数据)
- 临时编写简单脚本解决问题(如批量处理文本、解方程)
使用
code_interpreter时,只需要在 tools 里面加一项{ "type": "code_interpreter" },,这样chat_template会自动识别到该字段并输出一些使用信息,告诉模型如何给出代码,并告知模型这个代码可以被执行- 以 LongCat-Flash-Chat/blob/main/tokenizer_config.json 为例,其具体做法是先将
code_interpreter包装成一个类似function的格式,再统一输出,最终效果就是让模型知道可以调用code_interpreter执行代码("code"参数内容就是代码)
- 以 LongCat-Flash-Chat/blob/main/tokenizer_config.json 为例,其具体做法是先将
function 和 code_interpreter 整体对比
- 注:在实际应用中,两者常结合使用:
function处理外部交互,code_interpreter处理复杂计算,共同扩展大模型的能力边界维度 function(函数)code_interpreter(代码解释器)核心能力 调用预定义功能接口 动态执行代码逻辑 适用场景 外部服务调用、结构化任务 复杂计算、数据处理、脚本生成 调用格式 严格 JSON 格式指令 代码片段(如 Python) 灵活性 低(固定功能) 高(支持任意逻辑) 安全性 高 需沙箱隔离,风险较高
附录:chat-template 格式化
大部分开源模型的 chat-template 都是压缩为一行的,可读性较差
可以使用下面的代码重新存储 tokenizer 信息
1
tokenizer.save_pretrained(output_model_name)
这样会同步生成得到的
chat-template.jinja文件,整体格式是更可读的注:也可以使用大模型来帮忙格式化
附录:chat-template continue 语句的使用
- 老版本的
transformers中,调用tokenizer.apply_chat_template时 不支持chat-template中有continue语句 - 若遇到类似下面的错误时,升级
transformers版本后可以解决问题:1
jinja2.exceptions.TemplateSyntaxError: Encountered unknown tag 'continue'. Jinja was looking for the following tags: 'elif' or 'else' or 'endif'. The innermost block that needs to be closed is 'if'.
附录:Qwen2-72B-Instruct chat-template 使用示例
Qwen2-72B-Instruct 的 chat-template 非常简单
特别需要说明:当不增加 System Prompt 时, Qwen2-72B-Instruct 会默认将
"You are a helpful assistant."作为 System PromptQwen2-72B-Instruct/tokenizer_config.json 原始定义:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40{
"add_prefix_space": false,
"added_tokens_decoder": {
"151643": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151644": {
"content": "<|im_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151645": {
"content": "<|im_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
}
},
"additional_special_tokens": ["<|im_start|>", "<|im_end|>"],
"bos_token": null,
"chat_template": "{% for message in messages %}{% if loop.first and messages[0]['role'] != 'system' %}{{ '<|im_start|>system\nYou are a helpful assistant.<|im_end|>\n' }}{% endif %}{{'<|im_start|>' + message['role'] + '\n' + message['content'] + '<|im_end|>' + '\n'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant\n' }}{% endif %}",
"clean_up_tokenization_spaces": false,
"eos_token": "<|im_end|>",
"errors": "replace",
"model_max_length": 131072,
"pad_token": "<|endoftext|>",
"split_special_tokens": false,
"tokenizer_class": "Qwen2Tokenizer",
"unk_token": null
}- 注:Qwen2.5-72B-Instruct 的 chat-template 有调整,支持了 工具调用等,同时还修改了默认的 System Prompt 为
"You are Qwen, created by Alibaba Cloud. You are a helpful assistant."1
2
3
4
5{
...,
"chat_template": "{%- if tools %}\n {{- '<|im_start|>system\\n' }}\n {%- if messages[0]['role'] == 'system' %}\n {{- messages[0]['content'] }}\n {%- else %}\n {{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}\n {%- endif %}\n {{- \"\\n\\n# Tools\\n\\nYou may call one or more functions to assist with the user query.\\n\\nYou are provided with function signatures within <tools></tools> XML tags:\\n<tools>\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n</tools>\\n\\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\\n<tool_call>\\n{\\\"name\\\": <function-name>, \\\"arguments\\\": <args-json-object>}\\n</tool_call><|im_end|>\\n\" }}\n{%- else %}\n {%- if messages[0]['role'] == 'system' %}\n {{- '<|im_start|>system\\n' + messages[0]['content'] + '<|im_end|>\\n' }}\n {%- else %}\n {{- '<|im_start|>system\\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- for message in messages %}\n {%- if (message.role == \"user\") or (message.role == \"system\" and not loop.first) or (message.role == \"assistant\" and not message.tool_calls) %}\n {{- '<|im_start|>' + message.role + '\\n' + message.content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {{- '<|im_start|>' + message.role }}\n {%- if message.content %}\n {{- '\\n' + message.content }}\n {%- endif %}\n {%- for tool_call in message.tool_calls %}\n {%- if tool_call.function is defined %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {{- '\\n<tool_call>\\n{\"name\": \"' }}\n {{- tool_call.name }}\n {{- '\", \"arguments\": ' }}\n {{- tool_call.arguments | tojson }}\n {{- '}\\n</tool_call>' }}\n {%- endfor %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != \"tool\") %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n<tool_response>\\n' }}\n {{- message.content }}\n {{- '\\n</tool_response>' }}\n {%- if loop.last or (messages[loop.index0 + 1].role != \"tool\") %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n{%- endif %}\n",
...,
}
- 注:Qwen2.5-72B-Instruct 的 chat-template 有调整,支持了 工具调用等,同时还修改了默认的 System Prompt 为
Qwen2-72B chat-template 使用示例:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116model_name = "/Users/xxx/llm/model/Qwen2-72B-Instruct"
# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
messages = [
{
"role": "system",
"content": "${system_prompt}"
},
{
"role": "user",
"content": "${user_round_0}"
},
{
"role": "assistant",
"content": "${assistant_round_0}"
},
{
"role": "user",
"content": "${user_round_1}"
},
{
"role": "assistant",
"content": "${assistant_round_1}"
},
{
"role": "user",
"content": "${assistant_round_2}"
}
]
output = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(output)
# <|im_start|>system
# ${system_prompt}<|im_end|>
# <|im_start|>user
# ${user_round_0}<|im_end|>
# <|im_start|>assistant
# ${assistant_round_0}<|im_end|>
# <|im_start|>user
# ${user_round_1}<|im_end|>
# <|im_start|>assistant
# ${assistant_round_1}<|im_end|>
# <|im_start|>user
# ${assistant_round_2}<|im_end|>
# <|im_start|>assistant
messages = [
{
"role": "system",
"content": "${system_prompt}"
},
{
"role": "user",
"content": "${user_round_0}"
},
{
"role": "assistant",
"content": "${assistant_round_0}"
},
{
"role": "user",
"content": "${user_round_1}"
},
{
"role": "assistant",
"content": "${assistant_round_1}"
}
]
output = tokenizer.apply_chat_template(messages, add_generation_prompt=False, tokenize=False)
print(output)
# <|im_start|>system
# ${system_prompt}<|im_end|>
# <|im_start|>user
# ${user_round_0}<|im_end|>
# <|im_start|>assistant
# ${assistant_round_0}<|im_end|>
# <|im_start|>user
# ${user_round_1}<|im_end|>
# <|im_start|>assistant
# ${assistant_round_1}<|im_end|>
messages = [
{
"role": "user",
"content": "${user_round_0}"
},
{
"role": "assistant",
"content": "${assistant_round_0}"
},
{
"role": "user",
"content": "${user_round_1}"
},
{
"role": "assistant",
"content": "${assistant_round_1}"
}
]
output = tokenizer.apply_chat_template(messages, add_generation_prompt=False, tokenize=False)
print(output)
# <|im_start|>system
# You are a helpful assistant.<|im_end|>
# <|im_start|>user
# ${user_round_0}<|im_end|>
# <|im_start|>assistant
# ${assistant_round_0}<|im_end|>
# <|im_start|>user
# ${user_round_1}<|im_end|>
# <|im_start|>assistant
# ${assistant_round_1}<|im_end|>
补充:qwen_vl_utils.process_vision_info 函数的使用
专为 Qwen-VL(视觉语言)系列模型(如 Qwen-VL、Qwen2.5-VL)设计
- 主要作用是将多模态对话数据(包含图片或视频的文本结构)解析并读取为模型能够直接使用的 PIL Image 对象或 Torch Tensor(视频帧) ,同时整理视频处理的采样参数
函数签名与参数详解
1
2
3
4
5
6def process_vision_info(
conversations: Union[List[Dict[str, Any]], List[List[Dict[str, Any]]]],
return_video_kwargs: bool = False,
return_video_metadata: bool = False,
image_patch_size: int = 14,
) -> Tuple[...]- 参数说明:
参数 类型 说明 conversationsList[Dict]或List[List[Dict]]必填,多模态对话内容,支持单轮对话(单层列表)或多轮对话(嵌套列表),
每条字典中需包含role(角色)和content(内容),且content内部需含有image、image_url或video字段来指定视觉数据的路径或 URLreturn_video_kwargsbool默认为 False,
若为True,返回值中会多一个字段(即完整的视频处理参数字典video_kwargs)return_video_metadatabool默认为 False,
若为True,返回值video_inputs中会包含video_metadata信息image_patch_sizeint图像分块大小(Vision Transformer 的 Patch Size),通常与模型加载时的配置保持一致,影响图像缩放时的网格计算,
默认为14,是图片和视频共享参数 - 返回值说明:函数返回一个 三元组
(image_inputs, video_inputs, video_kwargs)或 二元组(image_inputs, video_inputs):返回值 类型 说明 image_inputsOptional[List[Image.Image]]读取后的 PIL Image 对象列表(按输入顺序),
如果输入中无图片,则为Nonevideo_inputsOptional[List[Union[torch.Tensor, List[Image.Image]]]]读取后的视频数据列表,
具体是torch.Tensor还是PIL Image列表取决于fetch_video的底层实现,通常为按帧采样的张量,如果输入中无视频,则为Nonevideo_kwargsOptional[Dict[str, Any]]视频采样辅助参数,
默认包含{'do_sample_frames': False},若满足特定条件,还会包含视频的fps(帧率)列表,用于后续模型处理时对齐时间维度,
仅return_video_kwargs=True时返回
- 参数说明:
内部处理流程
- 1)提取视觉信息 :调用
extract_vision_info(conversations)递归解析对话结构,提取出所有包含image、image_url或video的字典片段 - 2)遍历并读取数据 :
- 若字段为
image或image_url,调用fetch_image读取图片(支持本地路径或 HTTP URL),返回 PIL Image 并加入image_inputs - 若字段为
video,调用fetch_video读取视频(返回视频张量以及采样帧率video_sample_fps),并加入video_inputs - 若字段均不匹配,抛出
ValueError
- 若字段为
- 3)空值处理 :如果最终图片列表或视频列表为空,将其设为
None - 4)构造视频参数字典 :
- 基础字典为
{'do_sample_frames': False} - 关键兼容逻辑 :若
return_video_metadata为False(兼容 Qwen2.5-VL 旧版本),则向字典注入{'fps': video_sample_fps_list}
- 基础字典为
- 5)返回结果 :根据
return_video_kwargs决定是否返回video_kwargs(若为 False,该值仍会返回,但通常不会被使用)
使用示例
场景 1:仅处理单张图片(不用获取视频参数)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18from qwen_vl_utils import process_vision_info
# 构造单轮对话,包含一张本地图片
conversations = [
{
"role": "user",
"content": [
{"type": "image", "image": "/path/to/your/photo.jpg"},
{"type": "text", "text": "描述这张图片"}
]
}
]
# 不用获取视频参数
images, videos, video_kwargs = process_vision_info(conversations)
# images 此时为 [<PIL.Image>],videos 为 None
print(f"读取到 {len(images) if images else 0} 张图片")场景 2:同时处理图片和视频,并获取视频参数
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24conversations = [
{
"role": "user",
"content": [
{"type": "image", "image_url": "https://example.com/pic.png"}, # 支持URL
{"type": "video", "video": "/path/to/video.mp4"},
{"type": "text", "text": "对比视频和图片"}
]
}
]
# 开启返回视频参数
images, videos, video_kwargs = process_vision_info(
conversations,
return_video_kwargs=True,
return_video_metadata=False # 兼容旧版Qwen-VL
)
# 读取结果
if images:
print(f"图片数量: {len(images)}")
if videos:
print(f"视频张量列表: {videos}") # 通常为 [torch.Tensor]
print(f"视频采样参数: {video_kwargs}") # 包含 do_sample_frames 和 fps场景 3:多轮对话结构(自动展平提取所有视觉信息,无需手动处理嵌套)
1
2
3
4
5
6
7
8
9
10
11
12multi_turn_convs = [
[ # 第一轮
{"role": "user", "content": [{"image": "img1.jpg"}]},
{"role": "assistant", "content": [{"text": "这是第一张"}]}
],
[ # 第二轮
{"role": "user", "content": [{"image": "img2.jpg"}, {"video": "video.mp4"}]}
]
]
images, videos,_ = process_vision_info(multi_turn_convs)
# 函数会自动展平提取所有视觉信息,无需手动处理嵌套
Kimi-K3 (Kimi K3) 的 chat-template 说明
Kimi K3 没有使用 jinja 模版,而是调用 Python 文件直接实现编码,包含了多个 Python 文件和一个 BPE 模型文件
tiktoken.model(大约 2.5 M)编码示例:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34from transformers import AutoProcessor, AutoConfig
model_path = "/Users/jiahong/llm/model/Kimi-K3"
processor = AutoProcessor.from_pretrained(model_path, trust_remote_code=True)
# 读取本地图片
from PIL import Image
import numpy as np
image = Image.fromarray(np.random.randint(0, 255, (256, 256, 3), dtype=np.uint8))
messages = [
{
"role": "system",
"content": "You are a helpful assistant!",
},
{
"role": "user",
"content": [
{"type": "text", "text": "We have two images"},
{"type": "text", "text": "\n# image1:"},
{"type": "image", "image": image},
{"type": "text", "text": "\n# image2:"},
{"type": "image", "image": image},
{"type": "text", "text": "\nWhat animal is on the candy?"},
]
},
]
prompt = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=False,
thinking=True,
thinking_effort="high",
)
print(prompt)编码结果:
1
2
3
4
5<|open|>message role="system" type="thinking-effort"<|sep|>`thinking_effort` guides on how much to think in your thinking channel (not including the response channel), supported values include `low`, `medium`, `high`, and `max`.
Now the system is invoked with `thinking_effort=high`.<|close|>message<|sep|><|end_of_msg|><|open|>message role="system"<|sep|>You are a helpful assistant!<|close|>message<|sep|><|end_of_msg|><|open|>message role="user"<|sep|>We have two images
# image1:<|kimi_image_placeholder|>
# image2:<|kimi_image_placeholder|>
What animal is on the candy?<|close|>message<|sep|><|end_of_msg|><|open|>message role="assistant"<|sep|><|open|>think<|sep|>
附录:模型下载和导入问题
加载 tokenizer 时可能会报错:
1
ValueError: Error parsing line b'version https://git-lfs.github.com/spec/v1' in /Users/jiahong/llm/model/Kimi-K3/tiktoken.model
根因:
tiktoken.model在仓库里本是 Git LFS 托管文件首次(LFS 未拉取时)被读取过一次,tiktoken 的
read_file_cached把当时那个 LFS 指针文本按sha1(文件绝对路径)缓存到了系统临时目录,比如 MacOS 是下面的临时目录:1
/var/folders/.../T/data-gym-cache/d2ec580956c6e1aac12c4f0a8f1c09478d6a9c4f
即使后来 LFS 真实内容(2.7MB 的 BPE 词表,首行
IQ== 0)已下载到位,tiktoken 仍优先读这个过期的缓存指针 ,于是把version https://git-lfs.github.com/spec/v1当成token rank来解析而报错
关键点:缓存键是文件路径的 sha1 ,不校验文件内容/哈希(
load_tiktoken_bpe未传expected_hash),所以文件变了缓存不会自动失效解决方案
方案 A(推荐,根治) :删除该条过期缓存:
1
2
3
4import hashlib, tempfile, os
blobpath = "/Users/jiahong/llm/model/Kimi-K3/tiktoken.model"
cache_dir = os.path.join(tempfile.gettempdir(), "data-gym-cache")
os.remove(os.path.join(cache_dir, hashlib.sha1(blobpath.encode()).hexdigest()))- 或直接清空整个缓存目录:
rm -rf /var/folders/*/T/data-gym-cache(macOS)/rm -rf /tmp/data-gym-cache(Linux)
- 或直接清空整个缓存目录:
方案 B(临时绕过) :在
import tiktoken之前禁用缓存:1
2import os
os.environ["TIKTOKEN_CACHE_DIR"] = "" # 空字符串 → read_file_cached 直接读源文件亲测:使用方案 A,删掉过期缓存条目后没有出现问题