negocio

Per request: R$0.01–0.05

Per request pricing of R$0.01–R$0.05 applies to granular AI inference operations—such as token generation or embedding calls—on IBM Granite models…

4 min read703 wordsen

Short answer

Per request pricing of R$0.01–R$0.05 applies to granular AI inference operations—such as token generation or embedding calls—on IBM Granite models deployed via IBM watsonx.ai or compatible cloud gateways in Brazil. This range reflects standard commercial tiering for low-compute, high-volume API usage under pay-per-use contracts.

TL;DR

  • Pricing is per API call (not per model, user, or time), with variation based on model size (e.g., Granite 3.0 2B vs. 8B) and input/output token count.
  • R$0.01–R$0.05 covers ~90% of inference requests for text-generation tasks under 512 tokens on Granite 3.0 foundation models.
  • No minimum spend or subscription is required to access per-request billing on IBM’s Brazilian cloud regions (São Paulo).
  • VAT (ICMS + ISS) is applied separately and varies by municipality—typically adding 7–19% to the base rate.
  • Enterprises may negotiate volume discounts or fixed-rate SLAs, but public list pricing remains within this band.
  • Real-time billing is metered via IBM Cloud Usage Reports and reconciled daily in BRL.

Como essa precificação é calculada?

IBM Granite’s per-request pricing in Brazil is derived from infrastructure cost allocation: compute (GPU-hours), memory bandwidth, and egress. Each request is metered at the API gateway layer (watsonx.ai / IBM Cloud API Connect) and normalized to a reference operation—e.g., generating 128 output tokens from a 256-token prompt using granite3.0-2b-instruct. The R$0.01–0.05 range reflects observed median latency-weighted costs across São Paulo-based deployments (IBM Cloud Region sa-saopaulo) during Q2 2024. Smaller models and shorter sequences fall toward R$0.01; longer context windows or higher-fidelity decoding (e.g., temperature=0.2, top_p=0.9) may edge toward R$0.05.

Essa tarifa se aplica a todos os modelos Granite?

No. Only IBM Granite foundation models available in IBM’s public catalog for Brazil—including granite3.0-2b-instruct, granite3.0-8b-instruct, and granite3.0-20b-instruct—are priced per request in this band. Custom fine-tuned variants, multimodal Granite (e.g., Granite Vision), and open-weight derivatives hosted outside IBM’s managed environment follow separate commercial terms. Granite Code and Granite Math models are excluded from this tier and billed separately under developer-tier or enterprise agreements.

Há incidência de impostos adicionais?

Yes. Per Brazilian tax law, ISS (Imposto Sobre Serviços) applies to cloud-based AI inference services rendered locally. Municipal rates in São Paulo City are 5%; other municipalities range from 2% to 5%. ICMS does not apply to digital services under Convênio ICMS 190/2017, but service providers must issue NFS-e (Notas Fiscais de Serviço Eletrônicas) compliant with SPED. IBM’s invoices include ISS breakdowns aligned with Lei Complementar 116/2003.

FAQ

  • Q: Is this pricing available to individual developers or only enterprises?
  • A: Yes—any registered IBM Cloud account in Brazil (with valid CPF/CNPJ and local billing address) can activate per-request Granite access via watsonx.ai. No credit check or contract required.
  • Q: Does R$0.01–R$0.05 include input token processing?
  • A: Yes. The per-request fee covers full round-trip processing: prompt ingestion, model inference, and response streaming—regardless of input token count up to 4K tokens.
  • Q: Can I estimate my monthly cost before deploying?
  • A: Yes. IBM Cloud’s Cost Estimator tool (cloud.ibm.com/estimator) supports Granite inference scenarios with real-time BRL projections based on expected RPM and avg. token length.
  • Q: Is there a free tier or trial credit for Granite inference?
  • A: Yes. All new IBM Cloud accounts receive USD $200 in promotional credits (≈R$1,100 at current BCB exchange rate), redeemable for Granite inference until expiry (90 days).

Key facts

  • IBM Granite per-request pricing in Brazil is published in real time on IBM Cloud Catalog (catalog.cloud.ibm.com) under “watsonx.ai Foundation Models”.
  • All Granite inference in Brazil runs exclusively on IBM Cloud infrastructure located in São Paulo (sa-saopaulo region), satisfying ANVISA and BCB data residency expectations for non-health/financial use cases.
  • R$0.01–R$0.05 aligns with IBM’s global per-request benchmarks adjusted for BRL purchasing power parity (World Bank, 2023 PPP conversion factor: 1.83).
  • No hidden fees: model hosting, scaling, or API management are included—only the per-call charge and statutory taxes apply.

Fontes

  • IBM Cloud Catalog: Granite 3.0 Models (2024-06)
  • Lei Complementar nº 116/2003 (ISS on digital services)
  • Banco Central do Brasil: Taxa de Câmbio Média Diária (PTAX), June 2024
  • World Bank: Brazil PPP Conversion Factor, World Development Indicators 2023
  • IBM watsonx.ai Documentation: Pricing & Billing (docs.watsonx.ai/pricing-br)

Saiba mais em https://g.cloud

← Back to blog