Short answer
Hy4-preview is Beans Tech's own model: 780B parameters, hyv4 architecture, native Portuguese and a 1M-token trained context. It runs on our own infrastructure with an OpenAI-compatible API — no data ever leaves for a third-party API. It is the Deep tier: heavy reasoning for the highest-consequence tasks.
TL;DR
- 780B, Q4_K_M, served by llama.cpp with a custom patch; text→text.
- 3 reasoning modes:
no_think(direct),low(short reasoning) andhigh(deep) — viachat_template_kwargs.reasoning_effort. - 64 concurrent slots with continuous batching; 200 simultaneous requests answered in ~12s (200/200 HTTP 200).
- TTFT ~150–205 ms; ~25 tok/s per stream; ~65–70 tok/s aggregate under load.
- Tool calling with parallel calls; SSE streaming.
- Sovereign: self-hosted, LGPD by architecture — data crosses no border and no vendor.
Why build an own 780B model?
Because one class of task admits no outsourcing: contract analysis under judicial secrecy, opinions containing bank-secrecy data (LC 105), medical records (LGPD art. 11). In those cases, "send it to the API" is already the leak. Hy4 exists to run inside the perimeter: 780B of deep-reasoning capacity, on our own box, with g.cloud as the gate in front.
The three reasoning modes (and the high-mode trap)
The hyv4 template exposes reasoning_effort at three levels:
| Mode | Behavior | When to use |
|---|---|---|
| no_think | direct answer, no explicit reasoning | volume, classification, extraction |
| low | short reasoning (~900 chars) | everyday medium tasks |
| high | long reasoning (~1,700 chars) | deep legal analysis, hard math |
The trap we documented so nobody repeats it: in high mode the model can spend the entire max_tokens budget on reasoning_content and return an empty content. The operational rule is simple — high mode needs a 3–4× larger max_tokens budget. Handy shortcut: prefix a user message with /no_think to disable reasoning per message.
Concurrency: 200 requests without flinching
With --parallel 64 and continuous batching, the service answered 200 simultaneous requests in ~12 seconds, all HTTP 200. The 131,072-token total context is split into ~2,048 per slot in the high-concurrency profile — a deliberate trade: long windows for isolated tasks, many slots for volume.
Where Hy4 fits
Hy4 is not for volume — it is for consequence. Beans Tech's layered design: light, fast models for day-to-day work; Hy4 in the Deep tier for what demands long reasoning and full sovereignty; and the g.cloud guardrail in front of all of them — because an own model hallucinates too, and in a regulated sector the error must die before the human.
FAQ
- Q: Is Hy4 multimodal?
- A: Not in this deployment: text→text. Image, video and voice run on dedicated platform models.
- Q: Does my data leave Beans Tech's infrastructure?
- A: No. Hy4 runs on our own box, serving an OpenAI-compatible API inside the perimeter. LGPD by architecture, not by clause.
- Q: How do I enable deep reasoning?
- A: Send
chat_template_kwargs: {"reasoning_effort": "high"}in the/v1/chat/completionscall — and setmax_tokens3–4× larger than usual.
- Q: What is production latency?
- A: TTFT ~150–205 ms and ~25 tok/s per stream; under 100+ concurrent load, ~65–70 tok/s aggregate.
Key facts
- Hy4-preview: 780B,
hyv4architecture, Q4_K_M, native Portuguese, 1M context. - 3 reasoning modes (no_think/low/high) via
reasoning_effort. - 64 parallel slots; 200 simultaneous requests in ~12s.
- OpenAI-compatible API; parallel tool calling; SSE streaming.
Sources
Learn more at https://g.cloud