Skip to content
View bertholomus's full-sized avatar
  • Joined Aug 1, 2026

Block or report bertholomus

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
bertholomus/README.md

BertholomusAI

Reproducible distributed LLM inference, quantization research, and evidence-backed deployment recipes for NVIDIA GB10 / DGX Spark systems.

We publish the receipts that matter: pinned revisions, integrity manifests, exact runtime formulas, bounded capability gates, concurrency/context evidence, restart proof, limitations, licensing, and upstream credits.

Verified public work

Project Topology Focus GitHub Hugging Face
GLM-5.3-Flash NVFP4 4× GB10, TP4 1M-configured context, FP8 KV, DFlash2, tools, multimodal and long-context gates Recipe + evidence Evidence card
DeepSeek V4 Flash Graph-8 2× GB10 Reproducible two-node deployment and benchmarks Repository Model card
DeepSeek V4 Flash Graph-8 4× GB10, TP4 Four-node scaling, evidence and release recipe Repository Model card

Publication standard

  • Attribute upstream models, checkpoints, runtimes, kernels, and recipe lineage.
  • Separate configured capability from directly proven capability.
  • Publish machine-readable evidence and known limitations alongside headline results.
  • Never present third-party quantizations as BertholomusAI-created artifacts.
  • Mirror weights only when provenance and redistribution rights are clear.
  • Prefer exact revisions and hashes over mutable tags.

Current research direction

  • GLM-5.3 distributed inference on GB10 systems.
  • Full-download quantization workflows with explicit authorship and reproducibility.
  • Speculative decoding, long-context validation, and safe production supervision.
  • Hermes-native local-agent infrastructure.

Hugging Face: huggingface.co/bertholomus · X: @bertholomusai

Popular repositories Loading

  1. deepseek-v4-flash-0731-dspark-graph8-4xgb10 deepseek-v4-flash-0731-dspark-graph8-4xgb10 Public

    Verified four-GB10 DeepSeek-V4-Flash-0731 TP4 DSpark Graph-8 deployment and benchmarks

    Python 1

  2. deepseek-v4-flash-0731-dspark-graph8 deepseek-v4-flash-0731-dspark-graph8 Public

    Verified two-node GB10 DeepSeek-V4-Flash-0731 DSpark Graph-8 deployment and benchmarks

    Python

  3. hermes-agent hermes-agent Public

    Forked from NousResearch/hermes-agent

    The agent that grows with you

    Python

  4. glm-5.3-flash-nvfp4-gb10-tp4 glm-5.3-flash-nvfp4-gb10-tp4 Public

    Evidence-backed GLM-5.3-Flash NVFP4 TP4 deployment validation on 4x NVIDIA GB10.

    Python

  5. bertholomus bertholomus Public

    BertholomusAI — reproducible distributed LLM inference and evidence-backed GB10 deployment recipes.