Register and share your invite link to earn from video plays and referrals.

Box
@Box
Helping devs build the future of intelligent, content-driven apps using @Box. Check out our docs and sign-up for a free developer account on
3.1K Following    80.2K Followers
Muse Spark 1.3 is now available in Box AI. What our evaluation of Muse Spark 1.3 found: → 13% accuracy lead over comparably priced models on the Box Complex Work Eval → 42% faster than Muse Spark 1.2 using roughly a third fewer tokens → Strongest gains in Financial Services, Technology, and Energy
Show more
We evaluated Claude Opus 5.5 at @Box through our complex work evals. 3 things stood out compared to Opus 5: • More efficient: 30% faster, ~⅓ the tokens • Clearer: ~40% shorter answers • Accuracy gains in tech & finance This combination is critical as our customers scale AI across content-heavy workflows: less overhead to get the work done, and clearer analysis that’s easier to understand and act on when decisions matter. Excited to be a launch partner—coming soon to Box AI Studio Here is more on our evals and findings:
Show more
At Box, we've been testing Opus 5.5 on a variety of complex enterprise knowledge work tasks dealing with unstructured data with the Box Agent. Overall, we saw frontier capability levels, with major performance improvements over Opus 5. 63% fewer tokens used, 42% less verbosity, and 30% faster vs. Opus 5. And the model itself is cheaper, so this is a major win for any agentic computer use, coding, analytics, or data work that enterprises will be doing. Here are some examples of the task wins and performance gains across a variety of industry tests that we performed: • Financial services - due diligence (+39% task accuracy): A year of transaction records, with the job of finding every miscalculation in an acquisition target's pricing tool. Opus 5.5 scored a perfect result on every attempt in half the words Opus 5 used, consuming 82% fewer tokens overall. • Technology - cloud cost analysis (+65% task accuracy): Work out what a company should actually change about its cloud spend. Opus 5.5 picked the right basis for the retention calculation and kept the source data's unit conventions straight all the way through, so the number at the end actually holds up. It took half the time Opus 5 took, with 70% fewer tokens. • Consumer products - client account analysis (+17% task accuracy): Set the onboarding targets for a client account, reading across the signed contract, a satisfaction tracker and a team metrics sheet. The contract never states a senior/junior split, so Opus 5.5 derived it from the 18-person roster and showed the rule it used; several clients had a perfect 10 on individual survey questions, so it averaged each client's responses instead of crowning the single 10. It finished this one in half the time, on 78% fewer tokens. • Clinical diagnostics - data analysis (+15% task accuracy): Malaria rapid-test performance across a dry and a wet season: build the patient records out of two clinical PDFs, compute positive test rates by season and gender, and test whether parasite counts really differ between test-positive and test-negative patients. Opus 5.5 caught that the two groups' standard deviations differed more than 100-fold, re-ran it the right way, and found the dry-season difference didn't hold up after all. This accuracy gain came with a final answer that was half the length of Opus 5's, and also needed 78% fewer tokens end to end. Customers will be able to build AI Agents with Opus 5.5 shortly in the Box AI Studio.
Show more
Read our full evaluation:
🚀@digitalocean Managed Agents makes running long-running agents simple and handles third-party tool access through Action Gateway. Box is supported on Action Gateway, so you can run long-running, governed agents connected to Box to search, read, write, and act on your enterprise content.
Show more
DigitalOcean Managed Agents is now in public preview. Run Claude Code, Codex, or your own LangGraph agent in a runtime environment that pauses when idle. Put its tools behind one governed endpoint, and pick from 75+ open and proprietary models. One cloud, one bill. Prompts to get started available in the blog:
Show more
BoxWorks 2026 is where you get the inside track on where Box is headed next. → Keynotes from Box Co-founder and CEO Aaron Levie, @nvidia Founder and CEO Jensen Huang, and @intel CEO Lip-Bu Tan → A live walkthrough of Box's latest innovations → First look at the product roadmap from Box leadership What does your team need most from enterprise AI right now? Save your spot:
Show more
Box CEO @levie says that the more AI agents companies deploy, the more important the software powering them becomes. "Everybody always made this claim that once you have all these agents, then actually all they really need is a database underneath." "But the problem is that database has to be coordinated against possibly thousands of end users and then possibly tens of thousands of agents." "So the access controls, the business logic, being able to actually render the experience, the security, all these core systems are actually going to become more important in a world of all of these agentic workloads."
Show more
48 hours out. AI Trivia and Game Night before @WeAreDevs kicks off. Wed Sept 23 | 5-7pm | Guildhouse, San Jose Games, food, drinks, prizes. Hosted by @pagerduty, @box, @arizeai, and @EntireHQ Grab a spot:
Show more
New this year, join us for the first ever Dev Summit at BoxWorks. If you’re building agentic and content-driven enterprise applications and workflows you won’t want to miss this! 2 days of technical sessions and hands-on workshops covering everything from the latest in agentic sandboxes to file systems for agents. Speakers from @nvidia, @OpenAI, and Box and more coming soon! November 5-6 in San Francisco. Register here:
Show more
A claims system collects evidence into a case folder. Someone has to decide whether that evidence is complete and acceptable. That decision has to be recorded. This demo routes evidence files through Box Automate. Your app triggers the workflow, which can be fully customized. Box can assigns the approval task, extract metadata, send evidence for e-signature, and more. Box allows for a smooth process, and you define the process. Demo.👇
Show more
We put Grok 4.7 from @SpaceXAI to work on a $2 million insurance claim where one deductible error alone changes the calculation by $143,000. In this Box Agent preview, @grok reconciles the claim against the policy and supporting records. It catches an $82,000 duplicate invoice and a missing $64,000 supplier credit. It also explains why the Business Income waiting period doesn’t apply to Extra Expense. The result is a cited claims review for the adjuster, with final coverage and payment decisions left to the insurer. Explore Box AI Studio to build custom agents for your own document-heavy workflows.
Show more
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
0
76
1.5K
187
Forward to community
We built an incident triage workflow using Box and Jev, a new frontier model by @CompleteSkeptic optimized for decisions rather than text generation. Unstructured content in, typed decisions out.
Show more
Jev will be super helpful for agents to make split second decisions in workflows, data classification, judgment calls, and hundreds of other use-cases in the enterprise. Here's a quick demo with Box and Jev to make that real. The demo pulls an incident report from Box, asks whether it's customer-facing and how severe it is, moves the file into escalate, monitor, or review folders, and sets a metadata template instance with the result. This all happens nearly instantly and at almost no cost. You can imagine this in insurance claims, contract management, loan processing, security reviews, customer log analysis, and so on. Definitely a great new class of AI use-case.
Show more
🆕@typesafeai dropped Jev this week, a new frontier model optimized for decisions. We put it to work on an incident triage workflow in Box. Jev doesn't generate text. You send it state plus typed questions and it returns booleans with probabilities, enums, and scores in a single pass, so there's no output to parse and no schema to keep in line. The demo uses Jev to classify and route incidents. It pulls the incident report from Box, builds the state, asks whether it's customer-facing and how severe it is, moves the file into Escalate, Monitor, or Review folders, and sets a metadata template instance with the result. The answers comes back with a confidence estimate, so anything below threshold goes to Review for a human instead of routing itself. Check it out.👇
Show more
fx is a ~6 MB single binary with a 10µs cold start written in zig. Can run in a terminal, a CI sandbox, or the browser via WebAssembly. It's going to be fascinating to see how devs evolve their pipelines and processes when you can embed this kind of intelligence.
Show more
🚀@vercel's’ fx is a tiny open-source coding agent built to take action directly in a developer’s workspace. This demo extends that idea by giving fx access to the security policies, requirements, and release docs stored in Box. Using the Box CLI, fx retrieves an approved policy, compares it with local code, edits the files, adds missing test coverage, and prepares release evidence without being told which lines to change. It asks for human approval before uploading anything, then verifies every file after it lands in Box. Git stays where the code lives. Box stays where the approved docs and release proof live. fx connects the work across both. Demo below ↓
Show more
Claude’s new Slides experience can turn deal materials in Box into an editable credit committee briefing. In our demo, @claudeai reviews a borrower’s deal workspace, creates a five-slide presentation with financial visuals and source references, and flags conflicting information for review. We leave a comment directly on a slide, Claude revises it, and the finished deck is exported as PowerPoint and uploaded back to Box through MCP for the credit team to review. Connect Box to Claude to create and refine presentations using your enterprise content.
Show more
AI tools will keep changing, but content policy needs to travel with the content. See how Box Shield applies classification-based access controls to what AI agents can actually read, not just what they can download. 👇
Show more
🆕 @OpenAIDevs Agents API gives developers a hosted sandbox where agents run tools, coordinate work, and handle complex tasks. Box Mount brings enterprise content directly into that sandbox as files the agent can read, reason across, and produce new work from, using normal shell commands and file paths while Box Mount handles two-way sync automatically. In our demo, a lead agent mounts a deal room from Box, reads five source files, launches three specialist agents in parallel, and writes four reports back through Box Mount with an initial verdict. When a customer revises the MSA in Box, Box Mount syncs the new version into the same sandbox and the agents reassess automatically, refreshing all four reports while preserving the revised MSA as version two in Box. The result is a shared workspace where people and agents work on the same governed Box content without building custom file-transfer logic, with permissions, versions, and audit history preserved throughout. Box Mount is in private preview. Watch👇
Show more
OpenAI and Box are bringing your enterprise content directly into @ChatGPT. As the file system for AI, Box provides headless access to enterprise context in ChatGPT while keeping the content managed, governed and protected by your existing Box permissions. Demo ↓
Show more