Add Chinese model card and language navigation

#8
Files changed (2) hide show
  1. README.md +306 -0
  2. checksums.sha256 +1 -1
README.md CHANGED
@@ -19,8 +19,14 @@ tags:
19
  inference: false
20
  ---
21
 
 
 
 
 
22
  # <img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/logo.png" width="40" style="display: inline-block; vertical-align: middle; margin: 0;" alt="GroundAnything logo" /> GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
23
 
 
 
24
  **Model family:** [GroundAnything — DLM / parallel decoding](https://huggingface.co/GroundingPI/GroundAnything) · [GroundAnything-VLM — autoregressive](https://huggingface.co/GroundingPI/GroundAnything-VLM).
25
 
26
  **This repository contains the GroundAnything DLM checkpoint.** Use entropy-guided decoding for the main benchmark setting, or select optional self-speculative decoding.
@@ -308,3 +314,303 @@ The paper's cumulative infrastructure speedups compare execution implementations
308
  ## Ethical Considerations:
309
 
310
  Grounding predictions can miss small or occluded objects, repeat instances, or produce inaccurate text and coordinates. Validate localization quality on the intended task and inspect incomplete responses. GUI points describe image locations; the model does not execute interface actions. Perception outputs require task-specific validation before integration into physical systems.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
19
  inference: false
20
  ---
21
 
22
+ <!-- model-card-language: en -->
23
+
24
+ <a id="english"></a>
25
+
26
  # <img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/logo.png" width="40" style="display: inline-block; vertical-align: middle; margin: 0;" alt="GroundAnything logo" /> GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
27
 
28
+ **English** | [简体中文](#chinese)
29
+
30
  **Model family:** [GroundAnything — DLM / parallel decoding](https://huggingface.co/GroundingPI/GroundAnything) · [GroundAnything-VLM — autoregressive](https://huggingface.co/GroundingPI/GroundAnything-VLM).
31
 
32
  **This repository contains the GroundAnything DLM checkpoint.** Use entropy-guided decoding for the main benchmark setting, or select optional self-speculative decoding.
 
314
  ## Ethical Considerations:
315
 
316
  Grounding predictions can miss small or occluded objects, repeat instances, or produce inaccurate text and coordinates. Validate localization quality on the intended task and inspect incomplete responses. GUI points describe image locations; the model does not execute interface actions. Perception outputs require task-specific validation before integration into physical systems.
317
+
318
+ ---
319
+
320
+ <!-- model-card-language: zh-CN -->
321
+
322
+ <a id="chinese"></a>
323
+
324
+ # <img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/logo.png" width="40" style="display: inline-block; vertical-align: middle; margin: 0;" alt="GroundAnything 标志" /> GroundAnything:以极速并行解码实现精准视觉定位
325
+
326
+ [English](#english) | **简体中文**
327
+
328
+ **模型系列:** [GroundAnything — DLM / 并行解码](https://huggingface.co/GroundingPI/GroundAnything) · [GroundAnything-VLM — 自回归](https://huggingface.co/GroundingPI/GroundAnything-VLM)。
329
+
330
+ **本仓库提供 GroundAnything DLM 模型。** 主要基准评测采用熵引导解码,也可选择自推测解码。
331
+
332
+ <p align="center" style="margin: 0;"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/d926c66e5f0be4cad6f263fbef5cbcbe98102a42/assets/fig1-teaser.png" width="100%" alt="GroundAnything:广泛的视觉定位与并行视觉证据提取" style="display: block; margin: 0;" /></p>
333
+
334
+ ## 🔗 快速链接
335
+
336
+ - 🚀 **在线演示:** ZeroGPU 不支持 GroundAnything DLM 运行时。
337
+ - 🌐 **项目主页:** [GroundAnything](https://groundingpi.github.io/groundanything/)。
338
+ - 💻 **GitHub 代码:** [groundingpi/GroundAnything](https://github.com/groundingpi/GroundAnything)。
339
+ - 📄 **论文:** [arXiv:2609.39600](https://arxiv.org/abs/2609.39600)。
340
+
341
+ # 模型概述
342
+
343
+ ### 简介:
344
+
345
+ 自回归(AR)视觉定位模型将空间预测串行化,带来顺序生成的延迟,并为输出 token 施加因果顺序。我们将视觉定位视为视觉证据提取:物体、位置和空间关系共同受到图像与查询的约束,但它们之间的依赖关系并不意味着生成过程必须从左到右进行。这一区别使双向扩散成为自然的选择,让空间假设能够并行产生,并通过迭代去噪共同细化。
346
+
347
+ 我们提出 **GroundAnything**,一个拥有 4B 参数的视觉定位基础模型,通过分块去噪兼顾快速并行解码与精确定位。在 30 个视觉定位基准上,自回归版本 **GroundAnything-VLM** 以 **72.42%** 的综合表现刷新了同等规模模型的最佳水平,并与 GPT-6 Astra(**71.35%**)保持竞争力。采用**熵引导解码**的 GroundAnything 同样超越了这一规模下此前的最佳水平,平均达到 **61.75%**,高于基于 MTP 的快速模型 LocateAnything 的 **53.32%**。
348
+
349
+ 在论文报告的运行配置下,可选的**自推测模式**相较自回归版本实现了 **4.51× 加速**,同时 **COCO F1mIoU 下降 0.74 个百分点**。基础设施实验表明,逐步优化推理实现能够将并行解码转化为实际加速,为对延迟敏感的实际系统提供高效的视觉定位能力。
350
+
351
+ ### 演示视频
352
+
353
+ <video controls playsinline preload="none" width="100%" poster="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/demo-poster.jpg" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/demo.mp4"></video>
354
+
355
+ **并行解码**
356
+
357
+ <video controls playsinline preload="none" width="100%" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/decoding.mp4"></video>
358
+
359
+ ### 许可证与使用条款:
360
+
361
+ 我们的原创贡献采用 **Apache 2.0** 许可证,本项目不额外施加限制。第三方材料保留其适用的许可证,其中源自 Kimi 的材料及适用的衍生作品遵循 **Kimi K3 License**。使用或再分发组合后的模型包时,这些条款仍然适用。许可范围与完整许可文本请参阅 [LICENSE](LICENSE)。
362
+
363
+ ### 部署地域:
364
+
365
+ 全球。
366
+
367
+ ### 应用场景:
368
+
369
+ - 开放词汇目标定位与密集场景视觉定位。
370
+ - 指代表达理解与基于视觉提示的定位。
371
+ - 点定位、空间推理和 GUI 元素定位。
372
+ - 文本定位与 OCR,以及文档版面理解。
373
+ - 面向机器人、具身智能体和自主系统的感知研究。
374
+
375
+ ### 发布日志:
376
+
377
+ - **网页 [10/04/2026]:** 项目网页发布:[GroundAnything](https://groundingpi.github.io/groundanything/)。
378
+ - **代码 [10/03/2026]:** GitHub 代码发布:[GroundAnything](https://github.com/groundingpi/GroundAnything)。
379
+ - **论文 [09/30/2026]:** [GroundAnything](https://arxiv.org/abs/2609.39600)。
380
+
381
+
382
+ <details>
383
+ <summary>引用</summary>
384
+
385
+ ```bibtex
386
+ @misc{yu2026groundanythingreconcilingparalleldecoding,
387
+ title = {{GroundAnything}: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
388
+ author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
389
+ year = {2026},
390
+ eprint = {2609.39600},
391
+ archivePrefix = {arXiv},
392
+ primaryClass = {cs.CV},
393
+ url = {https://arxiv.org/abs/2609.39600},
394
+ }
395
+ ```
396
+
397
+ </details>
398
+
399
+ ## 模型架构:
400
+
401
+ **架构类型:** 共享视觉语言主干,支持自回归模型和分块扩散模型。
402
+
403
+ - **视觉编码器:** MoonViT-V2 / Kimi-K3 视觉主干。
404
+ - **语言主干:** Qwen3-4B-Instruct-2507。
405
+ - **多模态投影器:** 2 × 2 空间聚合与两层 MLP。
406
+ - **空间词表:** 1,000 个坐标 token,与语义标签和协议标记共用词表。
407
+ - **DLM 转换:** 共享解码器与词表输出头同时支持因果预测和双向响应块去噪,并为扩散生成加入掩码 token。
408
+
409
+ <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/fig2-architecture.png" width="100%" alt="GroundAnything:视觉语言架构与自回归到扩散的转换" /></p>
410
+
411
+ ## 输入:
412
+
413
+ **输入类型:** 图像与文本。
414
+
415
+ - **图像:** 每个请求包含一张 RGB 图像。Python 客户端支持 JPEG、PNG 和 WebP 文件;随模型提供的处理器负责图像预处理。
416
+ - **文本:** 自然语言指令、类别列表、指代表达、OCR/版面查询,或包含示例框的提示词。
417
+ - **HTTP API 的图像编码:** 在 `image_url` 内容项中提供 base64 data URI,并配合一个 `text` 内容项。
418
+
419
+ 请使用模型自带的分词器、处理器和对话模板。多个类别使用 `</c>` 分隔。参考框与输出使用相同的 0–999 空间 token 词表。
420
+
421
+ ## 输出:
422
+
423
+ **输出类型:** 包含语义标签和量化空间坐标的文本。
424
+
425
+ 三个已发布模型均采用 **GAM 协议**:使用 `<0>` 至 `<999>` 的整数坐标 token、目标引用分隔符和框分隔符。边界框包含 `(x1, y1, x2, y2)`,点包含 `(x, y)`。相邻坐标 token 之间没有空格。同一标签下的多个实例在同一组框分隔符内以逗号分隔。未找到的目标用 `None` 表示。
426
+
427
+ 语法示例:
428
+
429
+ ```text
430
+ <|object_ref_start|>car<|object_ref_end|><|box_start|><100><200><500><650><|box_end|>
431
+ <|object_ref_start|>car center<|object_ref_end|><|box_start|><300><425><|box_end|>
432
+ <|object_ref_start|>absent object<|object_ref_end|><|box_start|>None<|box_end|>
433
+ ```
434
+
435
+ 客户端返回解析后的预测结果,以及 `raw_output`、`finish_reason`、`usage`、`parse_error` 和 `valid` 字段,并将坐标映射回图像像素以便可视化。被截断或格式不正确的响应会标记为无效。自定义 API 客户端应使用 `skip_special_tokens=false` 保留空间 token,并避免在这些 token 之间插入空格。
436
+
437
+ ## 软件集成:
438
+
439
+ **运行引擎:** 仓库定制的 **SGLang** 集成支持 DLM 和 VLM 服务;此外还提供 DLM 的原生 Transformers 参考路径。
440
+
441
+ **默认服务环境:** Linux、Python 3.12,以及源码包提供的服务配置。使用 `python3 run.py setup serve` 安装随包提供的定制引擎及固定版本依赖;仅安装上游 SGLang 无法获得相同的模型适配与解码集成。
442
+
443
+ 服务配置固定使用 Torch **2.9.1**、Transformers **5.5.4**、Triton **3.5.1**、sgl-kernel **0.3.20** 和 FlashInfer **0.5.3**。服务与评测使用独立环境。
444
+
445
+ **已测试硬件:** NVIDIA **B300、B200、H200、H800**,以及 **PPU**。提供的 SGLang 服务安装器面向 CUDA/GPU;PPU 需要对应的平台运行时,该 GPU 安装器不会选择 PPU 运行环境。
446
+
447
+ 默认 DLM 服务采用 **BF16、Triton 注意力、eager 执行、单 GPU、一个活动请求和两个排队请求**。客户端并发会使请求进入队列,并不意味着模型以多请求批次运行。CUDA Graph 与选择性 FP8 作为独立的基础设施实验,在下文说明。
448
+
449
+ ## 模型版本:
450
+
451
+ | 模型 | 生成方式 | 评测模式 |
452
+ |:---|:---|:---|
453
+ | [GroundAnything](https://huggingface.co/GroundingPI/GroundAnything) | 熵引导分块扩散;可选自推测解码 | **GAM** |
454
+ | [GroundAnything-VLM](https://huggingface.co/GroundingPI/GroundAnything-VLM) | 自回归生成 | **GAM** |
455
+
456
+ GroundAnything 的主要结果使用**熵引导解码**;GroundAnything-VLM 使用**自回归解码**。自推测解码是一种可选加速模式。
457
+
458
+ ## 评测
459
+
460
+ 评测工具包支持 **7 种模式**:
461
+
462
+ | 模式 | 支持的模型 |
463
+ |:---|:---|
464
+ | **`GAM`** | **GroundingPI、GroundAnything、GroundAnything-VLM** |
465
+ | `VLM` | 通用视觉语言基线模型 |
466
+ | `REXOMNI` | Rex-Omni |
467
+ | `LOCATEANYTHING` | LocateAnything |
468
+ | `GROUNDINGDINO` | 通过兼容服务接入的 GroundingDINO |
469
+ | `DLM` | 使用 GAM 协议的旧版扩散模型 |
470
+ | `RLV2` | 使用 GAM 协议的旧版 RL 模型 |
471
+
472
+ **三个已发布模型均使用 GAM 模式。**
473
+
474
+ 评测代码与使用说明:[GitHub](https://github.com/groundingpi/GroundAnything)。
475
+
476
+ ## 定量评测
477
+
478
+ <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/fig7-grounding-performance.png" width="100%" alt="GroundAnything:GroundAnything 与 GroundAnything-VLM 的基准表现概览" /></p>
479
+
480
+ ## 推理:
481
+
482
+ ### 安装
483
+
484
+ 获取代码包后,请在 **GroundAnything 源码仓库根目录**中运行以下命令,环境为 **Linux x86_64 和 Python 3.12**。使用源码包提供的服务配置,以确保定制模型适配器、解码实现与依赖保持一致。
485
+
486
+ ```bash
487
+ python3 -m pip install -r requirements.txt huggingface_hub
488
+ python3 run.py setup serve
489
+ ```
490
+
491
+ 安装器会校验并解压随包提供的框架,创建 `.venv-serve`,并记录解析后的依赖包。它会准备定制的 SGLang 实现,并按照服务配置指定的依赖顺序完成安装。配置中包含必要的 cuDNN 兼容性选择;依赖检查报告会记录已知的 Torch/cuDNN 元数据例外。
492
+
493
+ **请根据所用模型选择下方对应的配置。** GroundAnything-VLM 用户只需执行 VLM 的下载和启动命令;DLM 命令需要另外下载 GroundAnything 权重。
494
+
495
+ ### GroundAnything:熵引导服务
496
+
497
+ 下载**完整的 DLM 模型包**,包括定制代码与分词器,然后启动服务:
498
+
499
+ ```bash
500
+ hf download GroundingPI/GroundAnything --local-dir weights/dlm_bundle
501
+ python3 run.py serve --decoder denoise
502
+ ```
503
+
504
+ 已发布的模型包本身就是 DLM 模型包,因此本次下载无需运行 `prepare-model`。服务地址为 `http://127.0.0.1:8101/v1`,模型 ID 为 `groundinganything`。
505
+
506
+ ### GroundAnything:自推测服务
507
+
508
+ 切换解码器前,先停止现有 DLM 服务:
509
+
510
+ ```bash
511
+ python3 run.py serve --decoder speculative
512
+ ```
513
+
514
+ 该模式复用相同的 DLM 权重和服务地址,采用扩散候选生成与贪心因果验证。
515
+
516
+ ### GroundAnything-VLM:自回归服务
517
+
518
+ 使用独立的自回归模型及其服务配置:
519
+
520
+ ```bash
521
+ hf download GroundingPI/GroundAnything-VLM --local-dir weights/vlm
522
+ python3 run.py serve --config configs/release/vlm_sglang.yaml
523
+ ```
524
+
525
+ 该服务使用 `http://127.0.0.1:8102/v1`,模型 ID 为 `groundinganything-vlm`。两个仓库共用空间输出接口,但模型加载和生成路径不同。
526
+
527
+ ### Worker(推荐)
528
+
529
+ 启动相应服务后,即可重复使用客户端:
530
+
531
+ ```python
532
+ from grounding_anything import GroundingAnything, visualize
533
+
534
+ client = GroundingAnything(
535
+ base_url="http://127.0.0.1:8101/v1",
536
+ model="groundinganything",
537
+ )
538
+ result = client.predict("example.jpg", "the red car", task="bbox")
539
+ print(result.to_dict())
540
+ if result.valid:
541
+ visualize("example.jpg", result).save("prediction.png")
542
+
543
+ point_result = client.predict(
544
+ "example.jpg", "the center of the red car", task="point"
545
+ )
546
+ ```
547
+
548
+ HTTP 客户端不持有模型权重。
549
+
550
+ ### 支持的任务与提示词模板
551
+
552
+ | 任务 | 提示词示例 |
553
+ |:---|:---|
554
+ | 类别 / 密集目标定位 | `Locate all the instances that match the following categories: car</c>person.` |
555
+ | 指代表达框定位 | `Locate the target referred to by the following description: the red car.` |
556
+ | 类别点定位 | `Point to: car</c>person.` |
557
+ | 指代表达点定位 | `Point to the target referred to by the following description: the red car.` |
558
+ | OCR | `OCR task detect all the text in box format.` |
559
+ | 版面 | `Detect all document layout elements that match the following categories: title</c>text.` |
560
+ | GUI | `Point to the UI element to click for the following instruction: open the settings menu.` |
561
+ | 视觉提示 | 以原生空间 token 格式提供参考框,再请求定位相似物体。 |
562
+
563
+ ```text
564
+ Given reference boxes <|box_start|><100><200><500><650><|box_end|> indicating one or more objects, find all similar objects in the image and output their bounding boxes.
565
+ ```
566
+
567
+ 便捷客户端的 `predict()` 方法封装了指代表达框定位和点定位提示词。其他任务模板请通过兼容 OpenAI 的请求调用 `/chat/completions`;图像与提示词应放在同一条用户消息中。
568
+
569
+ ### 生成模式
570
+
571
+ #### 熵引导解码
572
+
573
+ 图像与查询的预填充会建立因果前缀缓存和一个已知锚点。发布配置采用**块大小 32、子块大小 4、熵阈值 0.8**。一个物理块包含已知锚点和 31 个掩码位置;子块从左到右依次完成,而每次前向计算覆盖整个物理块。
574
+
575
+ 对于当前子块中仍被掩码覆盖的每个位置,解码器基于可生成词表上未经修改的 token 分布计算熵,其中不包含掩码 token。熵不高于阈值的位置会被同时确定。如果没有位置满足阈值,则确定熵最低的位置,以保证解码继续推进。已确定的 token 保持不变。
576
+
577
+ 一个块完成后,通过一次因果前向计算重建该块的正式 KV 缓存,并提供下一个锚点。这次缓存构建**不会验证或拒绝已生成的块**。因此,一个需要 D 次去噪的块会执行 D + 1 次模型前向计算,不计最初的预填充。
578
+
579
+ #### 自推测解码
580
+
581
+ 模型使用自身共享的权重,通过双向注意力生成候选 token,再通过因果注意力进行验证。验证接受**最长的连续匹配前缀**,在首次不匹配处停止,应用因果分支的修正,并丢弃被拒绝后缀的缓存状态。发布的推测路径采用贪心验证,并非通用的随机推测采样器。
582
+
583
+ <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/fig6-self-speculative-decoding.png" width="100%" alt="GroundAnything:共享模型权重的线性与二次自推测调度" /></p>
584
+
585
+ 文档中的 `--decoder speculative` 服务采用线性的共享权重路径。精确贪心验证是相对于转换后模型的因果分支而言,并不意味着其输出与单独训练的 GroundAnything-VLM 模型相同。
586
+
587
+ ## 推理基础设施
588
+
589
+ ### SGLang 执行
590
+
591
+ 定制的 SGLang 集成负责协调模型加载、请求调度、注意力内核,以及分块生成过程中的 KV 缓存管理。去噪在当前块内部使用双向注意力,已完成的历史内容则保留为因果 KV。自推测解码还会验证候选并移除被拒绝的后缀状态。优化执行过程时,必须保留这些缓存语义。
592
+
593
+ 提供的配置选择 **BF16 + Triton 注意力 + eager 执行**。DLM 配置将引擎来源、解码器、实际生效设置和依赖包版本记录在 `outputs/sglang/<decoder>/engine_runtime.json` 中。VLM 服务则在 `outputs/sglang/vlm/engine_runtime.json` 中单独记录其因果模型与运行时配置。如需提高服务并发量,可部署使用不同设备、端口和输出目录的独立副本;默认队列并非多请求连续批处理。
594
+
595
+ ### CUDA Graph 与选择性 FP8
596
+
597
+ 论文评估了逐步叠加的基础设施优化:
598
+
599
+ | 层级 | 目的 | 提供的默认配置中的状态 |
600
+ |:---|:---|:---|
601
+ | 原生 PyTorch eager | 模型执行参考路径 | 参考实现 |
602
+ | SGLang eager | 集成调度、注意力与缓存执行 | **默认服务路径** |
603
+ | CUDA Graph 重放 | 重放已捕获且兼容的 GPU 操作,减少重复启动开销 | 论文已评估;**默认启动器禁用** |
604
+ | 选择性 FP8 | 降低适用的语言模型线性运算的计算成本,其他组件保留 BF16 | 论文已评估;**默认服务仍使用 BF16** |
605
+
606
+ CUDA Graph 改变的是兼容 GPU 操作的提交方式,不会定义新的 token 确定或验证规则。捕获的形状与状态更新必须兼容当前解码路径。选择性 FP8 可能改变 logits 以及后续的解码决策,因此属于不同的数值配置。
607
+
608
+ 可选的图实现会捕获固定形状的块计算。其 FlashInfer 路径采用持久化注意力掩码和设备端缓冲区更新,支持双向候选生成与因果验证。Triton 验证器路径将候选生成与验证的元数据和输入缓冲区分开,同时共享参数与实际 KV 池。支持的捕获场景为 B32、单请求,且不扩展张量、流水线或数据并行;预填充和不支持的形状采用 eager 执行。可选的影子检查会将缓存状态、logits 和 token 选择与 eager 执行结果进行对比。本次发布没有通过公开的 `run.py serve --cuda-graph` 开关提供这些实现路径。
609
+
610
+ 论文中的基础设施累计加速是在**同一种解码模式内**比较不同执行实现得到的,与自推测解码相对于自回归模型的主要加速对比相互独立。上方默认发布命令不会启用 Graph 重放或 FP8,也不表示存在未提供的启用参数。
611
+
612
+ ## 伦理考量:
613
+
614
+ 视觉定位预测可能遗漏较小或被遮挡的目标、重复输出实例,或产生不准确的文本与坐标。请针对目标任务验证定位质量,并检查不完整的响应。GUI 点表示图像中的位置,模型本身不会执行界面操作。感知输出在集成到物理系统前,需要进行针对具体任务的验证。
615
+
616
+ <!-- /model-card-languages -->
checksums.sha256 CHANGED
@@ -1,5 +1,5 @@
1
  c419c191a33774f4e1f834b3fb36033dc84e8ba316a3d1935eed1fa529884a74 LICENSE
2
- 27a4f178ec745e626f5d3ffd540078e38ffcca7977b17c837d2849308accda5a README.md
3
  669bb095cca2e86ddc821926f4c3b42dd7389f2ad3ec49c8edca8fe6294aae7f added_tokens.json
4
  a0bc6f6fc7a29a80017a433e8f03a1cc1236e838a944a2d034295a60c4f2fddb chat_template.jinja
5
  22c369e2bdac19723fe7db7e4c64b224462744b1f1aae7bd5daf05307db3b9fd config.json
 
1
  c419c191a33774f4e1f834b3fb36033dc84e8ba316a3d1935eed1fa529884a74 LICENSE
2
+ 3189f325228af1f532159028f03b81c1ad871e163e47737a9e70e1d68ed3772d README.md
3
  669bb095cca2e86ddc821926f4c3b42dd7389f2ad3ec49c8edca8fe6294aae7f added_tokens.json
4
  a0bc6f6fc7a29a80017a433e8f03a1cc1236e838a944a2d034295a60c4f2fddb chat_template.jinja
5
  22c369e2bdac19723fe7db7e4c64b224462744b1f1aae7bd5daf05307db3b9fd config.json