今天凌晨 3 点,阿里开源发布了新推理模型 QwQ-32B,其参数量为 320 亿,但性能足以比肩 6710 亿参数的 DeepSeek-R1 满血版。


开源地址:

https://modelscope.cn/models/Qwen/QwQ-32B
https://huggingface.co/Qwen/QwQ-32B

Qwen Chat免费体验:

https://chat.qwen.ai/?models=Qwen2.5-Plus


模型效果

QwQ-32B 在一系列基准测试中进行了评估,包括数学推理、编程和通用能力。以下结果展示了 QwQ-32B 与其他领先模型的性能对比,包括 DeepSeek-R1-Distilled-Qwen-32B、DeepSeek-R1-Distilled-Llama-70B、o1-mini 以及原始的 DeepSeek-R1。

可以看到,QwQ-32B 的表现非常出色,在 LiveBench、IFEval 和 BFCL 基准上甚至略微超过了 DeepSeek-R1-671B。


强化学习
QwQ-32B 的大规模强化学习是在冷启动的基础上开展的。
在初始阶段,先特别针对数学和编程任务进行 RL 训练。与依赖传统的奖励模型(reward model)不同,千问团队通过校验生成答案的正确性来为数学问题提供反馈,并通过代码执行服务器评估生成的代码是否成功通过测试用例来提供代码的反馈。
随着训练轮次的推进,QwQ-32B 在这两个领域中的性能持续提升。

在第一阶段的 RL 过后,他们又增加了另一个针对通用能力的 RL。此阶段使用通用奖励模型和一些基于规则的验证器进行训练。结果发现,通过少量步骤的通用 RL,可以提升其他通用能力,同时在数学和编程任务上的性能没有显著下降。


API
如果你想通过 API 使用 QwQ-32B,可以参考以下代码示例:
from openai import OpenAI
import os

# Initialize OpenAI client
client = OpenAI(
    # If the environment variable is not configured, replace with your API Key: api_key="sk-xxx"    
    # How to get an API Key:https://help.aliyun.com/zh/model-studio/developer-reference/get-api-key    
    api_key=os.getenv("DASHSCOPE_API_KEY"),    
    base_url="https://dashscope.aliyuncs.com/compatible-mode/v1"
)

reasoning_content = ""
content = ""

is_answering = False

completion = client.chat.completions.create(
    model="qwq-32b",    
    messages=[
            {"role": "user", "content": "Which is larger, 9.9 or 9.11?"}    
            ],    
            stream=True,    
            # Uncomment the following line to return token usage in the last chunk    
            # stream_options={    
            #     "include_usage": True    
            # }
)
print("\n" + "=" * 20 + "reasoning content" + "=" * 20 + "\n")

for chunk in completion:    
      # If chunk.choices is empty, print usage
      if not chunk.choices:        
          print("\nUsage:")        
          print(chunk.usage)    
      else:        
        delta = chunk.choices[0].delta
        # Print reasoning content        
        if hasattr(delta, 'reasoning_content') and delta.reasoning_content is not None:            
           print(delta.reasoning_content, end='', flush=True)            
           reasoning_content += delta.reasoning_content        
        else:            
           if delta.content != "" and is_answering is False:                
              print("\n" + "=" * 20 + "content" + "=" * 20 + "\n")                
              is_answering = True            
        # Print content            
        print(delta.content, end='', flush=True)            
        content += delta.conten


消费级显卡即可本地部署

千问QwQ-32B既能提供极强的推理能力,又能满足更低的资源消耗需求,非常适合快速响应或对数据安全要求高的应用场景,开发者和企业可以在消费级硬件上轻松将其部署到本地设备中,进一步打造高度定制化的AI解决方案。
此外,千问QwQ-32B模型中还集成了与智能体Agent相关的能力,使其能够在使用工具的同时进行批判性思考,并根据环境反馈调整推理过程。通义团队表示,未来将继续探索将智能体与强化学习的集成,以实现长时推理,探索更高智能进而最终实现AGI的目标。