mercado

AntAngelMed: the 103B medical MoE with 6B active

AntAngelMed is Ant Healthcare's open-source medical model with Zhejiang's health information center: a 103B MoE with only 6.1B active per inference,…

3 min read573 wordsen

Short answer

AntAngelMed is an open-source medical language model developed by Ant Healthcare with Zhejiang's Provincial Health Information Center (China). It is a MoE with 103B total parameters and only 6.1B active per inference — which makes it fast (>200 tokens/s) and cheap to operate. Apache-2.0 license, 128K context.

TL;DR

  • MoE architecture based on Ling-flash-2.0 (inclusionAI), 1/32 activation ratio.
  • #1 overall on MedBench v4, leading in 5 dimensions; best open-source on HealthBench (with a strong margin on the Hard subset).
  • 103B total / 6.1B active → large-model quality at small-model inference cost.
  • 128K context (YaRN); >200 tok/s on H20; official FP8 and community GGUF builds available.
  • 3-stage training: medical continued pretraining → heterogeneous SFT (math, code, clinical dialogue) → RL with GRPO and dedicated reward models.
  • Released 2025-12-12 (Hugging Face: MedAIBase/AntAngelMed).

Why MoE changes the medical-model math

The dilemma of hosting medical AI: big models are good enough and too expensive; small models fit the budget and make mistakes. AntAngelMed breaks the dilemma with Mixture-of-Experts: 103B of stored knowledge, but only 6.1B activated per token. In practice, 100B-class quality at 6B-class throughput — over 200 tokens/s on a single H20. For a hospital or healthtech, that means clinical triage, record summarization and decision support running on your own infrastructure, with no patient data leaving the perimeter.

What the benchmarks say

On MedBench v4 (the most rigorous Chinese medical benchmark, physician-evaluated), AntAngelMed ranks #1 overall, leading in 5 dimensions. On OpenAI's HealthBench it is the best open-source model, with a strong margin on the Hard subset — precisely the difficult clinical cases. Its declared strength is medical Q&A and ethics/safety, a dimension other models tend to neglect.

The guardrail is still required

Topping a medical benchmark authorizes no one to practice medicine. In Brazil: diagnosis requires a CFM-registered physician (Law 12,842/2013), telemedicine follows CFM Resolution 2,314/2022, medical-purpose software falls under ANVISA's RDC 657/2022, and medical records are sensitive personal data (LGPD art. 11). g.cloud's job is to make sure AntAngelMed's output — or any model's — passes through the gate before reaching a human: no diagnosis without a physician, no sensitive data leaking, a public receipt for every decision.

FAQ

  • Q: Is AntAngelMed free for commercial use?
  • A: Yes, Apache-2.0. Open weights, modification and self-hosting allowed.
  • Q: What hardware does it need?
  • A: Full BF16: 8× Ascend 910B (64 GB) or 4× Kunlun P800/PPU 810 (96 GB). INT4: 2× Ascend 910B. An official FP8 build and community GGUF for llama.cpp exist.
  • Q: Does it work in English or Portuguese?
  • A: Training is centered on Chinese and English. For clinical use in other jurisdictions, validate first — and keep the scope guardrail and citation verifier active.
  • Q: How do I integrate it with g.cloud?
  • A: AntAngelMed serves via vLLM/SGLang with an OpenAI-compatible API; g.cloud plugs in as proxy, SDK or gateway plugin, without changing the model.

Key facts

  • AntAngelMed: MoE 103B total / 6.1B active, Apache-2.0, 128K context.
  • #1 on MedBench v4; best open-source on HealthBench (strong on Hard).
  • >200 tokens/s on H20; official FP8; community GGUF (mradermacher).
  • Built by Ant Healthcare + Zhejiang Provincial Health Information Center + Zhejiang Anzhen'er Medical AI.

Sources

Learn more at https://g.cloud

← Back to blog