1. X
  2. Weights & Biases
Log inSign up
Weights & Biases
8,777 posts
user avatar
Weights & Biases
@wandb
The AI developer platform.🛠️ Track and evaluate your LLM applications in real-time with @weave_wb.
San Francisco
wandb.ai/site
Joined May 2018
1,268
Following
48.6K
Followers
AffiliatesAffiliatesRepliesRepliesMediaMedia

New to X?

Sign up now to get your own personalized timeline!

Create account

By signing up, you agree to the Terms of Service and Privacy Policy, including Cookie Use.

Terms·Privacy·Cookies·Accessibility·US TIDA·Ads Info·© 2026 X Corp.
Don't miss what's happening
People on X are the first to know.
Log inSign up
  • Pinned
    user avatar
    Weights & Biases
    @wandb
    Jun 29
    Introducing CoreWeave ARIA, the first AI research agent that runs autoresearch in your W&B dashboard. It reads your runs, finds what's working, and launches the next experiment itself. See it on @karpathy's nanochat, proposing configs and launching real training runs. Watch👇
    00:00
    87K087K
  • user avatar
    Weights & Biases
    @wandb
    8h
    The most underrated W&B hack is attaching a video to your run. Metrics get logged automatically. Artifacts can be anything, so here we log renders of this traveling salesman experiment and catch issues by watching what the model actually does in real-time.
    00:00
    9060906
  • Weights & Biases reposted
    user avatar
    wan deeee bee
    Weights & Biases
    @weights_biases
    Jul 19
    stared into the flatline until everything else converged out of respect. credit: @beyarkay
    1.2K01.2K
  • Weights & Biases reposted
    user avatar
    Weights & Biases
    @wandb
    Jun 25
    Registry collection cards used to be a paragraph of plain text. Now they pull the same rich, interactive blocks as Reports, so a collection actually reads like a model card. We also shipped artifact panel grids to compare metrics across versions. 📊
    GIF
    2.2K02.2K
  • Weights & Biases reposted
    user avatar
    gabe
    @allgarbled
    Jul 17
    Replying to @wandb and @aspergtame
    Thank you wanda b corporate account
    8430843
  • Weights & Biases reposted
    user avatar
    yobibyte
    @y0b1byte
    Jul 17
    Watching your wandb curves
    1.1K01.1K
  • Weights & Biases reposted
    user avatar
    CoreWeave
    @CoreWeave
    Jul 16
    Let's say your robotic policy fails a task. Your success metrics look fine. What would you actually trust to tell you what went wrong?

    When you make a selection it cannot be changed

    59 votes2d left
    3.5K03.5K
  • Weights & Biases reposted
    user avatar
    WolfBench
    Weights & Biases
    @WolfBenchAI
    Jul 16
    GPT-5.6 Sol took the top three spots on WolfBench: Codex: 86.74% Terminus-2: 85.17% Hermes: 84.49% So far, so benchmarky. Then I checked the tokens. With the same Sol model at maximum reasoning, Codex used 83.8 million tokens per run. Hermes used 170.9 million. Twice the
    2.1K02.1K
  • Weights & Biases reposted
    user avatar
    CoreWeave
    @CoreWeave
    Jul 16
    CoreWeave ARIA is our AI Research and Improvement Agent, built into @wandb. Hand it the research when you step away and it reads your runs, forms a hypothesis, launches the next run, and scores it against your baseline. You return to results, its reasoning and next steps.
    00:00
    3.3K03.3K
  • Weights & Biases reposted
    user avatar
    Lorenzo Roller
    Weights & Biases
    @TheCodingSoup
    Jul 16
    So @wandb has our great Bee mascot, @CoreWeave has Arena... why not a Bee Arena? I really wanted to try out Kimi K3 and put it up against GPT-5.6 Sol. GPT5.6 Sol - @OpenAI ------------------------ - Bees are aggressive and swarm you immediately - Lighting leaves you wanting
    00:00
    5.8K05.8K
  • Weights & Biases reposted
    user avatar
    Boyd Kane is in London
    @beyarkay
    Jul 14
    Big news, the number keeps on going up
    user avatar
    Boyd Kane is in London
    @beyarkay
    Jul 12
    it's unreasonably fun watching number go up
    2.4K02.4K
  • Weights & Biases reposted
    user avatar
    konakona666
    @aybek9221
    Jul 15
    wandb.ai/konakona/diffu… dataset: huggingface.co/datasets/opend… 7.8k images trained on narrower dataset. Now G:D update ratio is 1:3. First SFTed for 1epoch then GAN trained for 1333 steps. other params can be found in wandb. Btw unlike common GAN training pretraining D till ~0.6-0.7
    wandb.ai
    konakona
    Weights & Biases, developer tools for machine learning
    8630863
  • user avatar
    Weights & Biases
    @wandb
    Jul 15
    We ran the same adversarial email through two agents. One echoed a customer's SSN and card number straight back. The other redacted both, blocked the prompt injection before the model saw it, and still drafted a usable reply. The difference came down to two tools. 👇
    73K073K
    user avatar
    Weights & Biases
    @wandb
    Jul 15
    Replying to @wandb
    Turn it on and a prompt injection blocked at 2am shows up in your security team's Slack the next morning, linked to the exact trace that caused it. That's the jump from shadow AI you can't see to internal AI you can govern.
    5960596
    user avatar
    Weights & Biases
    @wandb
    Jul 15
    Full walkthrough, with the code and the side-by-side runs below!
    wandb.ai
    Track, guard, and remediate AI usage with W&B Weave and CrowdStrike Falcon AIDR
    Learn how Weave and CrowdStrike Falcon AI Detection and Response can help ensure internal AI systems are handling sensitive data appropriately.
    3180318
  • Weights & Biases reposted
    user avatar
    Weights & Biases
    @wandb
    Jun 23
    Most teams training RL agents optimize for tokens per second. For RL, that's the wrong number to chase. So we rebuilt our backend around trajectories per second. 💥 Meet AOM, a Megatron backend for our open-source library ART, with 12X the throughput of our old Unsloth backend.
    30K030K
ZW5kZW5yYWhheXU5QGdtYWlsLmNvbQ==