arquitetura

Hy4-preview 780B: sovereign deep reasoning, no third-party API

Hy4-preview is Beans Tech's own 780B model: native Portuguese, 1M-token trained context, three reasoning modes (no_think/low/high), an OpenAI-compatible…

3 min read590 wordsen

Short answer

Hy4-preview is Beans Tech's own model: 780B parameters, hyv4 architecture, native Portuguese and a 1M-token trained context. It runs on our own infrastructure with an OpenAI-compatible API — no data ever leaves for a third-party API. It is the Deep tier: heavy reasoning for the highest-consequence tasks.

TL;DR

  • 780B, Q4_K_M, served by llama.cpp with a custom patch; text→text.
  • 3 reasoning modes: no_think (direct), low (short reasoning) and high (deep) — via chat_template_kwargs.reasoning_effort.
  • 64 concurrent slots with continuous batching; 200 simultaneous requests answered in ~12s (200/200 HTTP 200).
  • TTFT ~150–205 ms; ~25 tok/s per stream; ~65–70 tok/s aggregate under load.
  • Tool calling with parallel calls; SSE streaming.
  • Sovereign: self-hosted, LGPD by architecture — data crosses no border and no vendor.

Why build an own 780B model?

Because one class of task admits no outsourcing: contract analysis under judicial secrecy, opinions containing bank-secrecy data (LC 105), medical records (LGPD art. 11). In those cases, "send it to the API" is already the leak. Hy4 exists to run inside the perimeter: 780B of deep-reasoning capacity, on our own box, with g.cloud as the gate in front.

The three reasoning modes (and the high-mode trap)

The hyv4 template exposes reasoning_effort at three levels:

| Mode | Behavior | When to use |

|---|---|---|

| no_think | direct answer, no explicit reasoning | volume, classification, extraction |

| low | short reasoning (~900 chars) | everyday medium tasks |

| high | long reasoning (~1,700 chars) | deep legal analysis, hard math |

The trap we documented so nobody repeats it: in high mode the model can spend the entire max_tokens budget on reasoning_content and return an empty content. The operational rule is simple — high mode needs a 3–4× larger max_tokens budget. Handy shortcut: prefix a user message with /no_think to disable reasoning per message.

Concurrency: 200 requests without flinching

With --parallel 64 and continuous batching, the service answered 200 simultaneous requests in ~12 seconds, all HTTP 200. The 131,072-token total context is split into ~2,048 per slot in the high-concurrency profile — a deliberate trade: long windows for isolated tasks, many slots for volume.

Where Hy4 fits

Hy4 is not for volume — it is for consequence. Beans Tech's layered design: light, fast models for day-to-day work; Hy4 in the Deep tier for what demands long reasoning and full sovereignty; and the g.cloud guardrail in front of all of them — because an own model hallucinates too, and in a regulated sector the error must die before the human.

FAQ

  • Q: Is Hy4 multimodal?
  • A: Not in this deployment: text→text. Image, video and voice run on dedicated platform models.
  • Q: Does my data leave Beans Tech's infrastructure?
  • A: No. Hy4 runs on our own box, serving an OpenAI-compatible API inside the perimeter. LGPD by architecture, not by clause.
  • Q: How do I enable deep reasoning?
  • A: Send chat_template_kwargs: {"reasoning_effort": "high"} in the /v1/chat/completions call — and set max_tokens 3–4× larger than usual.
  • Q: What is production latency?
  • A: TTFT ~150–205 ms and ~25 tok/s per stream; under 100+ concurrent load, ~65–70 tok/s aggregate.

Key facts

  • Hy4-preview: 780B, hyv4 architecture, Q4_K_M, native Portuguese, 1M context.
  • 3 reasoning modes (no_think/low/high) via reasoning_effort.
  • 64 parallel slots; 200 simultaneous requests in ~12s.
  • OpenAI-compatible API; parallel tool calling; SSE streaming.

Sources

Learn more at https://g.cloud

← Back to blog