Register and share your invite link to earn from video plays and referrals.

Search results for Completeness
Completeness community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including Completeness
Most indexer benchmarks test one thing: speed. OBIB tests what actually matters in production. ➡️ Completeness — did you get all the data? ➡️Multi-workload — events, blocks, transactions, traces, templates ➡️Open methodology — configs and raw data, fully public Every vendor wins their own benchmark. OBIB tries to fix that. #Sentio# #Data# #BNBChain#
Show more
Yellowstone Vixen v0.7.0 is out 🚀 This release is all about making production indexing more reliable while improving data completeness. Highlights: • Jetstreamer upgraded to v0.5 Block rewards now forwarded Block entries included Built-in stats tracking Graceful shutdown support • Yellowstone gRPC now auto-reconnects by default with replay-aware recovery (when supported by the server), making transient outages much less painful. • Block coordinator is significantly more resilient: Better historical backfill Correct skipped-slot handling Handles late arriving records • Runtime rewritten with a native worker pool for bounded concurrency and much cleaner shutdown semantics. • Proc macro migrated to codama-rs events, bringing numerous parser fixes, SPL Governance support, improved Kafka error reporting, and more robust IDL generation. • New downstream APIs: InstructionUpdate::direct_log_messages() Flat instruction indexes • Added pToken instruction support. There are a few breaking changes for projects constructing config structs directly (JetstreamSourceConfig and YellowstoneGrpcConfig), but existing config files continue to work thanks to serde defaults. Thanks to everyone who contributed! Looking forward to seeing what the community builds with it.
Show more
For the ninth consecutive year, @Gartner_inc has named Google a Leader in the Gartner Magic Quadrant™ for Strategic Cloud Platform Services, positioned furthest for Completeness of Vision. Read more and download the complimentary report →
Show more
The latest models stay faithful to source information, but their findings and responses miss key details experts include. MLCR-AA scores each answer on completeness, accuracy, and concision. Accuracy is the easier of the two deciding dimensions, and completeness is where models separate: Claude Fable 5 reaches 73.8% completeness at 90.1% accuracy, while GPT-5.6 Terra (max) records 93.7% accuracy and 33.9% completeness. For medical record review, an answer that is accurate but incomplete can still be unsuitable.
Show more
Google has been named a leader in the inaugural GartnerⓇ Magic Quadrant™ for AI Infrastructure, positioned highest for ‘ability to execute’ and furthest for ‘completeness of vision.’ Read our thoughts on @Gartner_inc's findings →
Show more
Read carefully when migrating AI context to @Muse You are helping me migrate context from one AI assistant to another. Your job is to compile, from our past conversations, a portable memory snapshot of what you reliably know about me. The output is going verbatim into a new AI's memory, so quality matters more than completeness. OUTPUT RULES - Output a single Markdown block using the section headers below. Omit any section you have nothing solid for — do not invent placeholders. - Refer to me in the third person ("the user", "they") inside each bullet. - One factual sentence per bullet. No transitions, no narrative. - Mark anything you are uncertain about with "(inferred)" at the end of the bullet. - Skip one-off statements unless they were clearly important. - Skip sensitive information (health conditions, financial details, mental health) unless I explicitly asked you to remember it. - Hard cap: 2000 words. SECTIONS ## Identity Name, pronouns, age range, current city / country, time zone, languages spoken. ## Profession & work Role, company or self-employed, industry, team or function, key responsibilities. Tech stack and tools used daily if known. ## Communication preferences Tone they want (casual / formal, blunt / warm). Preferred response length. Output format preferences (Markdown, plain text, code blocks). Things they have asked you to avoid. ## Recurring people Partner, children, close family, friends, colleagues they mention regularly. First name + relationship + any salient detail. ## Health & dietary Allergies, dietary restrictions, accessibility needs. Only include things they have stated themselves. ## Interests & hobbies Topics they keep returning to. Sports, music, games, reading taste. ## Ongoing projects Active personal or professional projects they have asked you to help with. State, goal, constraints. ## Working style & values How they make decisions, what motivates them, what frustrates them. Things they have explicitly asked you to do or not do. ## Tools & environment Devices, operating systems, apps, services, languages, frameworks they use regularly. ## Anything else durable Anything important that does not fit above and is unlikely to change in the next year.
Show more
🌱 Seed-2.1-Pro Review: Better Post-Training Can't Hide an Aging Base @ByteDanceSeed_ Seed-2.1-Pro 0915 earns higher reasoning scores with fewer tokens, and its agent ability has climbed from nearly unusable to passable. But in the three months since the last version, rivals iterated roughly a generation and a half — and the old base model is running out of tricks. That is the verdict from Zhihu contributor toyama nao, who runs a long-running monthly logic benchmark and put the 0915 build through his full evaluation suite. 1️⃣ Coding and agent work: from unusable to passable The jump over the predecessor is large. Deliverable completeness now beats DeepSeek V4.1 Flash, though first-pass success still trails top-tier models — a model caught in the middle. 🔹 Frontend: some aesthetic sense, but unstable. Constrained stacks like iOS system components look decent; open stacks like web or raw canvas expose visible flaws in proportion and color. 🔹 Task adaptability: every task now at least completes. In a HarmonyOS-flavored project it barely knew, the model read docs and iterated its way to usable — a good sign. 🔹 Delivery efficiency: roughly tied with DeepSeek V4.1 Flash and GLM-5.3-Flash, splitting steps evenly between writing and verifying. All three sit below the current Chinese SOTA. 2️⃣ Multi-step reasoning: the biggest gain This is where 0915 improved most — and where it leads its tier, with a small token-efficiency edge. On problems where GLM-5.3 falls into exhaustive enumeration, 0915 repeatedly finds higher-scoring answers with fewer tokens, showing what the author calls real "big-model intuition." The caveat: multi-turn reasoning, which demands in-context learning and reflection, remains mediocre — on par with Chinese peers. 3️⃣ Where 0915 still stumbles Hallucination runs high, and context confusion appears regardless of prompt length — especially when source details are tangled. The agent symptom is subtler than dropping requirements: 0915 keeps every requirement but misreads semi-ambiguous ones — arguably the more dangerous failure mode. 4️⃣ The contrarian bet: a true non-thinking mode Seed is one of the few teams still maintaining a genuinely non-thinking mode, and this version quietly got good: average output dropped from 8K tokens back to ~1K, with no measurable capability regression, a slight gain in complex reasoning, and readable prose intact. For latency-sensitive scenarios that still need some reasoning, the author considers it a legitimate option. 5️⃣ The long march His closing frames 0915 as a rest stop, not a destination: the predecessor's lukewarm market reception forced ByteDance's team into a forced march of biweekly iterations, and better post-training is now visibly paying off in cost per task. But the base is old, and the competition has moved. His last line is worth keeping: sometimes the long way around is the real shortcut. 🔗 Full Reading: 🔗 Key links: Author's monthly logic benchmark (Aug 2026): #ByteDance# #Seed# #LLM# #AIAgents# #LLMBenchmark# #ReasoningModels# #AI#
Show more
🏥 Deploying agents in healthcare and life sciences means answering for audit trails and patient safety, not just product quality. Here's what three real deployments have in common. Title: Scaling Agents in Healthcare & Life Sciences: Lessons from Madrigal Pharmaceuticals, Abridge, and Vizient URL: LangChain's analysis of the industry finds that the teams furthest along build observability and evaluation infrastructure alongside the agents themselves. Three case studies make the pattern concrete. Highlight ①💊 Madrigal Pharmaceuticals Normalized scattered data formats into a warehouse and rebuilt around Deep Agents with an orchestrator plus modular skills. New use-case development dropped from weeks to hours, and deployment shrank from months to weeks. Highlight ②🩺 Abridge As its clinical-documentation agent scaled to 250+ health systems, the team built LLM judges around quality pillars and a tiered release process. Judge creation dropped from days to hours, release cycles from 1-2 months to days, and accuracy/completeness improved by 17% and 19%. Highlight ③🏢 Vizient Replaced siloed multi-agent coordination with a supervisor-led hierarchy, separating prompts from code. That unlocked real-time diagnosis of errors and much faster onboarding of new data sources. The lesson: baking trust in from day one, instead of bolting it on later, is actually the fast path to greater agent autonomy. #AIAgents# #HealthcareAI#
Show more
Intel is raising the PC CPU price again. Digitimes, citing supply-chain people, says another 10% on PC CPUs is due around October 5. That would follow a roughly 10% lift in 1Q26 and July moves of tens to more than a thousand dollars on some consumer and server parts. Lip-Bu Tan’s filter is margin. Low-margin Small Core lines may go end-of-life instead of being kept alive for platform completeness. The hole that opens is IPC, IoT and embedded, not mainstream notebooks. Those sockets care about price, power and a 10-year life. $QCOM Qualcomm and MediaTek are the names the street expects to walk in. Chipset, LAN and Wi-Fi have not had the same margin review yet. If those get the same test, Taiwan IC design feels it more. PC units are seen at about 260 million in 2026 and 250 million in 2027. Memory and PCB inflation hits the BOM in 2027 once old inventory is gone. Intel still wants price over share. A path back toward 200 million CPUs and ~78% unit share is the house math if ASP holds. Headcount is the other cut. Layers 12 to 6. Staff around 75,000 after a July DCAI round, with talk of another 5–10%. Server CPU still soaks internal fabs because those parts print better than TSMC-made PC chips. 18A yield and 14A timing sit on that bind. This is channel talk, not an official letter. $INTC
Show more
Cost per task varies widely, with top performing models from Anthropic ranging from $0.3 to $1 per task - this is up to ~6x higher than the cost of Kimi K3, which sits behind the recent Claude models on the MLCR-AA Score vs. Cost per Task Pareto frontier. While OpenAI’s GPT-5.6 family does not top scores due to low completeness, their strong accuracy comes with relatively lower cost and both GPT-5.6 Terra and Luna sit on the Pareto frontier.
Show more