Register and share your invite link to earn from video plays and referrals.

Search results for cowcat
cowcat community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including cowcat
A fun fact about Vanuatu is that they named their provinces by concatenating the first letters or syllables of the islands within them.
AGENTIC TRAFFIC NOW MAKES UP MORE THAN 70% OF ALL INFERENCE TRAFFIC 🚀 Agentic workloads are characterized by four elements: 🟠 Multi-turn: a session includes tens or hundreds of turns, leading to high potential KV-cache reuse. 🟠 Long context: system prompts, tool definitions, and the large number of turns make context accumulate quickly. 🟠 High prefix reuse: since the conversation progresses linearly, where output from turn n-1 is concatenated to turn n (typically), most context can be served from KV cache rather than recomputed (this depends on the amount of storage available to store KV tensors). As n grows, the ratio of cached input relative to uncached input typically tends towards 1. 🟠 Sub-agent bursts: a session launches multiple short-lived sub-agents with fresh context, which create bursty KV-cache patterns.
Show more
0
54
615
104
Forward to community
Introducing FlyOCR 🪰 - I trained a fly brain to read a PDF It uses the full MaleCNS v1.0 fruit fly connectcome. The architecture is inspired by doomfly by @wormuth The fly splits a pdf image into individual glyphs, maps pixels into receptor activations, runs simplified current-based dynamics across the 166k neurons and 25m edges in the circuit, applies a compact readout model on the downstream spikes, and concatenates everything into the parsed output. On reading an actual Microsoft 10-k, the fly gets ~86% over the balance sheet heading, but is largely able to read the numeric values correctly. Over 1.7k+ sampled glyphs (chars+digits) it gets 87% accuracy. With enough training it might match some of the latter-generation MNIST models! Maybe eventually we’ll replace our doc parsing VLMs with flies. Full video below. Repo with full code + report:
Show more
fyi we had around ~10k new launches with the new pre-IPO feature release. just imagine what would've happened if we didn't enforce ticker locks + client limits on new launches. we would've probably seen 100k launches within a single day. this is where a platform becomes a product used by bots rather than real humans, who can't possibly filter through so many new assets. this ends up driving FOMO trades and chasing volume, which is not hard to inflate. now also imagine what might be the effect of these launch limits constantly being tweaked with more demand coming in, esp once LONG will start scale very fast(hint concatenation of liquidity + flows into the best and most unique assets in the eco) we are going to win so much that you'll get tired of winning. next steps: 1. I want to apply filters that help users discover assets based on a variety of params (whale %, age, anti-fragility and some surprises). I think this is something we can innovate on as well. just better discovery. 2. we are actively monitoring any suspicious activity across all pairs. as I said, it's not hard to tweak our params. I don't see anything critical as of now, but we're already preparing some new params we can execute on with a click of a button, as soon as we feel like PNDs are becoming coordinated rather than organic market activity(remember AI nuked by 80%+ like 4 times) 3. As always, I think the best trading style on LONG is to find a team/project you think can cook and grow out of bad periods, help them out, and build a nice Longfolio of PVE assets. Keep in mind all of our top pairs were able to break ATHs or sustain value weeks post launch. LONG.
Show more
🚨 On September 6, 2026, @Liquid_BTC was affected by a cache key collision vulnerability in rangeproof verification. An attacker minted ~3,998.5 L-BTC with no corresponding peg-in. Within minutes, the unbacked L-BTC was pegged out into real BTC on the Bitcoin mainnet. About 3,400 BTC was later returned to the federation peg wallet, while ~598.5 BTC remains under the attacker’s control. The SlowMist Security Team traced the fund flows on the #Bitcoin# side using @MistTrack_io and fully analyzed the incident. 🧩 Attack flow: 1️⃣ Two setup transactions first landed valid rangeproofs and commitments, while embedding a crafted payload in the locking script to seed node caches. 2️⃣ A follow-up minting output reused a colliding cache key — the same raw concatenation of proof, commitment, asset commitment, and scriptPubKey, but with different field boundaries. 3️⃣ On a cache hit, nodes skipped secp256k1_rangeproof_verify and min-value checks, accepted an unbacked commitment, and minted ~3,998.5 L-BTC. The fake UTXOs were consolidated and pegged out within minutes. ⚙️ Root Cause: The Elements rangeproof cache key concatenated variable-length fields without length prefixes. Distinct argument tuples could hash to the same key, so a positive cache hit meant skipping cryptographic verification. 🛡️ SlowMist Insight: A positive-result cache in a consensus verification path is itself a cryptographic primitive. Every field the verifier reads — and every field boundary — must be unambiguously bound into the key. Treat cache-key integrity as a mandatory item in consensus-layer audits. Full analysis👇
Show more
# Codex Features and Practical Usage 📝 Stop repeating the same context in every prompt. AGENTS.md is a layered instruction file that teaches Codex your build steps and code conventions for good. 🏷️ Title: Hierarchical Project Instructions (AGENTS.md) 🔗 URL: 📘 Overview AGENTS.md files are persistent instructions Codex reads before starting work. Encode your project norms, build commands, and code standards in Markdown, and Codex automatically discovers and applies them. The defining trait is a layered model where files closer to your current directory take precedence. ⚙️ How It Works Codex builds an instruction chain with strict precedence. ・Global (`~/.codex/`): checks `AGENTS.override.md` first, then `AGENTS.md` ・Project (Git root down to cwd): at each level checks `AGENTS.override.md`, then `AGENTS.md`, then custom fallback names ・Merge: files concatenate from the root downward with blank lines; later files (closer to cwd) override earlier ones It stops adding files once the combined size hits `project_doc_max_bytes` (32 KiB default). Naming conventions: ・`AGENTS.md`: the standard instruction file ・`AGENTS.override.md`: a temporary replacement that takes precedence at its level ・fallback names: configurable via `project_doc_fallback_filenames` (e.g. `TEAM_GUIDE.md`) 🛠️ Practical Usage Split by role for best effect. ・`~/.codex/AGENTS.md` (global): universal personal preferences, e.g. "Always run `npm test` after editing JS," "Prefer `pnpm`" ・repo-root `AGENTS.md`: team norms like linting standards, documentation expectations, and review criteria ・subdirectory `AGENTS.override.md` (e.g. `services/payments/`): override broader rules for a specific domain without deleting parent guidance Verify what loaded by asking Codex to summarize the current instructions, e.g. `codex --ask-for-approval never "Summarize the current instructions."` 💡 Use Cases In a monorepo where only the payments service needs a different test command or review bar, dropping a `services/payments/AGENTS.override.md` switches the rules on only when Codex works under that path. The "Review guidelines" you write here also apply to GitHub's `@/codex review`. ⚠️ Caveats Codex rebuilds the instruction chain on every run, so there is no cache to clear — if guidance looks stale, restart in the target directory. Empty files are skipped and non-existent fallback names are ignored. The `CODEX_HOME` environment variable overrides the default profile location. #OpenAICodex# #AGENTSmd#
Show more
Fable 5.1 Masked a green screen on multiple shots And it did it around 90% quality of a human! Instruct your Claude: 1. Prompt a flat, unobstructed green screen. 2. Get word timings for slide cues. 3. Detect green per frame (HSV threshold). 4. Restrict detection near previous quad. 5. Fit four-corner quad to blob. 6. Smooth quads over five frames. 7. Mask only green inside quad. 8. Never key green outside quad. 9. Foreground objects preserved by mask. 10. Cover-fit slide, crop evenly. 11. Perspective-warp slide into quad. 12. Hard-cut slides at cue timestamps. 13. Despill: clamp G to max(R,B). 14. No add-back, avoids magenta. 15. Green-hued pixels get full alpha. 16. Raise thresholds where hands cross. 17. Static shot: freeze reference quad. 18. Fast motion: unsmoothed per-frame fit. 19. Video slide: warp every frame. 20. Grade slide: darker, slight blur. 21. Re-render only changed clips. 22. Normalize clips before stitching. 23. Stitch with concat demuxer only. 24. Verify audio per segment. 25. Check frames with your eyes. If you still think AGI won't happen this year I invite you to think again After effects people Your days are numbered
Show more
# Weaviate Features and Practical Usage 🚀 Ever wished you could delete all the embedding-API plumbing from your app code? Weaviate's model provider integrations let you wire up vectorization, generation, and reranking just by writing it into your collection config. 📌 Title and Feature URL Title: Model provider integrations URL: 📝 Overview Weaviate integrates with 20+ model providers including OpenAI, Cohere, Google, AWS, Azure OpenAI, Mistral, Anthropic, Hugging Face, and Ollama. You can plug them into automatic embedding at import, automatic embedding of query text, generation for RAG, and reranking of search results. The big win is that your application no longer needs code to call an embedding API and pass vectors in. 🔧 How It Works Integrations fall into three roles: - Vectorizer (embeddings): text or multimodal vectorization. - Generative (LLM): text generation for RAG pipelines. - Reranker: result refinement (offered by Cohere, Jina AI, NVIDIA, Voyage AI). There are also two delivery forms. API-based providers (OpenAI, Google, Cohere, AWS Bedrock, etc.) call external services, while locally hosted options (Ollama, Hugging Face Transformers, Model2vec) run on your own infrastructure. API-based modules are enabled by default in v1.33+. 🛠 Practical Usage - Specify an embedding provider with Configure.Vectors at collection creation, and Weaviate vectorizes automatically at both import and query time. - Configure generation with Configure.Generative to run RAG over your search results. - Configure a reranker with Configure.Reranker. - Automatic vectorization targets text / text[] properties. Weaviate sorts property names alphabetically, concatenates them, optionally prepends the collection name, and sends the string to the model (you can also exclude properties per-field). 🎯 Use Cases - Internal document search: auto-vectorize body text at import, and auto-vectorize the query with the same model so they stay consistent. - Model swapping: change vendor or model by editing collection config only. - Closed-network requirements: use a locally hosted option like Ollama to keep data in-house through embedding generation. - RAG chat: combine retrieval and generation within the same configuration, minimizing external orchestration. ⚠️ Caveats - API-based providers require API keys and incur usage charges. - Rate limits follow each provider's policy; watch out during bulk imports. - For versions before v1.27, the concatenated string is lowercased before being sent to the model. - For versions before v1.33, set ENABLE_API_BASED_MODULES to use API-based modules. #Weaviate# #Embeddings#
Show more
Tons of interesting things in the Kimi K3 tech report — here are five algorithm-side techniques that I think either I've never seen before or simply deserve more attention than they're getting. 1/ They open-sourced the model but kept the speculative decoding draft model, which could be a real serving advantage of their own. Quick background for those less familiar — speculative decoding pairs a large target model with a small draft model. The draft cheaply proposes several tokens ahead, and the target verifies them all in one forward pass. The speedup is determined by the acceptance rate, i.e., how often the target agrees with the draft's proposals. K3 is pre-trained with an MTP (multi-token prediction) layer, DeepSeek-V3 style — an extra layer on top of the backbone that predicts one token further into the future than the main next-token head. Structurally this layer is an exact copy of a regular backbone block, so you can think of it as a 94th layer that has been trained on the full pre-training corpus from day one, just with a shifted prediction target. After post-training, they freeze the target and fine-tune this MTP layer into an EAGLE-3-style draft. The EAGLE family of methods makes the draft a single decoder layer that reads the target model's internal hidden features rather than only the generated token sequence — conditioning on the target's features is what lets a one-layer draft stay accurate. The fine-tuning is then set up to match inference exactly. At inference, the draft proposes multiple tokens in a row, so from the second token onward it is building on its own unverified guesses rather than anything the target has confirmed. They replicate this condition during training by unrolling the draft for 7 steps — the first step uses the target model's features, and every step after that consumes the draft's own outputs from earlier steps. Two more design choices worth knowing. The draft reads low/mid/high-level target features (outputs of the 1st, 4th, and final AttnRes blocks), concatenated and passed through a fusion matrix initialized as [0 0 I] — zero weights on the low and mid features, identity on the high-level one. At initialization the draft therefore sees exactly the high-level feature the MTP layer was pre-trained on, and it gradually learns to mix in the other two during fine-tuning. And instead of the usual KL surrogate, they directly minimize the negative log of the acceptance rate itself (the sum of min(p, q) over the vocabulary, where p and q are the target and draft distributions), since minimizing KL does not guarantee maximizing acceptance for a capacity-limited draft. Everything is trained under the same MXFP4/MXFP8 QAT as their serving stack. The release itself is asymmetric. The full target weights are on HuggingFace, but the draft — and as far as I can tell, the MTP layer it was fine-tuned from — is not. That MTP layer was trained jointly with the backbone on the full pre-training corpus, which no external party has access to. So first-party serving stays faster and cheaper on the exact same open weights. This is the smartest business decision I've seen recently from open-weight model companies. 2/ Sync RL with partial rollouts. Sync RL waits for every rollout in the batch to finish before updating, and since rollout lengths vary wildly, compute is wasted by waiting on all rollouts to complete. Async RL decouples actors from the learner, which is much more efficient, but actor weights go stale, and a long rollout can land several learner steps behind the current policy. K3 runs a middle ground, where generation pauses as soon as a fraction of the trajectories completes, and optimization proceeds immediately, like async. However, unfinished trajectories get paused, enqueued, and resumed at the start of the next iteration under the freshly updated policy. In other words, a single 1M-token trajectory can literally be a relay across several different policy versions. Essentially, they trade model staleness for data staleness — off-policy prefixes inside otherwise on-policy trajectories — and mitigate it with a per-token regularization that constrains each update to a localized neighborhood of the current policy. It's an intellectually pleasing trade. What makes this viable at 1M context is the environment side, since pausing the model's rollout is easy; but pausing a live agentic sandbox mid-trajectory is not. Their microVM runtime checkpoints an environment in 133ms and resumes it in 49ms, and a paused sandbox consumes zero CPU and memory. 3/ Multi-teacher on-policy distillation as the merge step. After RL, they have nine expert models — three domains (general / agents / coding) crossed with three reasoning-effort levels (low / high / max) — consolidated into one unified model through multi-teacher OPD. The idea itself is not new; the report cites the same lineage as Thinking Machines' OPD post, MiMo-V2-Flash, and DeepSeek-V4. Three details stand out though. First, this is distillation with zero compression. Teachers and student are the same 2.8T architecture, and OPD is purely the mechanism that folds nine RL policies into one model, not a way to shrink a big teacher into a small student. Second, the OPD signal is implemented as a per-token RL reward — the clipped log-ratio between the teacher's and the student's probability of each generated token. Distillation is therefore not a separate pipeline; it is literally the same RL trainer running with a different reward. The student generates its own on-policy rollouts, the teacher scores every token along the way, and everything above carries over for free — partial rollouts, pausable sandboxes, the per-token regularization — which is what makes it feasible to distill even million-token agentic trajectories. Third, a negative result. At each step the student samples one token from its own distribution and it is all the OPD reward looks at, requiring just a single number per step, the teacher's log-prob of that token. They experimented with finer-grained top-k objectives that match more of the teacher's distribution over candidate tokens at each step, and saw no advantage in either convergence speed or final performance. So they decided that no full logits were needed. 4/ The RL harness is randomized. They represent an unified agent harness with shared tool interfaces, system prompts, context management strategies, skills, memories, subagents, and can instantiate Kimi Code, Claude Code, Codex, OpenClaw, Hermes, or entirely new harnesses from the same abstraction. During RL, harness configurations are dynamically reshuffled across task groups so the model never overfits to any single tool schema or interaction protocol. The implication is that harness generalization is a trained property, not an emergent one. If you have ever evaluated open models across different agent scaffolds and wondered why some transfer well and some fall apart, this is probably a big part of the answer. It also fits Kimi's position as an open-weight company. A closed lab ships the model and the harness together and controls the whole stack; an open model gets dropped into whatever scaffold people already use — Claude Code, Codex, OpenClaw, some custom internal agent, so harness robustness is even more importnat for open-weights models. Interestingly, their own in-house coding bench even reports K3 scoring slightly higher under Claude Code than under their own Kimi Code. 5/ NoPE on every global attention layer. All 24 Gated MLA layers in K3 use no positional encoding at all. Positional and recency information is carried entirely by the KDA layers' gating and decay (the backbone runs 3 KDA per 1 MLA), while the MLA layers do pure content-based global lookup. The payoff shows up at context extension. K3 grows from 8K to 64K during pre-training and from 256K to 1M during cooldown with zero positional-encoding modification — no RoPE base retuning, no YaRN. Hybrid linear attention is usually pitched as the efficiency component of these architectures; here the linear layers are also doing the entire job of the position encoding. Honestly, this only scratches the surface. The infra sections (MoonEP, quantile balancing, KDA-aware prefix caching) each deserve a post of their own. Full report is definitely worth the read.
Show more
# Neo4j Features and Practical Usage 🔍 A query that becomes a five-way JOIN in SQL can be written in a single line of arrow-drawing — Cypher is the declarative query language for graphs. 🏷️ Title: Cypher (Declarative Graph Query Language) 🔗 URL: 📘 Overview Cypher is a declarative graph query language designed for Neo4j. Instead of specifying "how to retrieve," you express "what data you want" using ASCII-art-like patterns. It is the common language of graphs, shared across application, analytics, and operations. ⚙️ How It Works The core of Cypher is pattern matching. ・Nodes are written in parentheses `(node)` ・Relationships use square brackets and arrows `-[rel]->`, with direction shown by the arrow ・Labels, types, and properties can be written directly inside the pattern (`(:User {id: $id})`) The main clauses are: ・`MATCH`: search for patterns within the graph ・`RETURN`: project the results to return ・`WHERE`: filter with conditions ・`CREATE` / `MERGE`: create / create-if-not-exists-else-match ・`WITH`: compose multi-step pipelines Cypher aligns with GQL (the ISO graph query language standard) while keeping Neo4j extensions. As of Neo4j 2025.06, new features are added exclusively to Cypher 25, while Cypher 5 is frozen. 🛠️ Practical Usage A query that becomes multiple JOINs in SQL can be written as a single pattern. ```cypher MATCH (c:Customer {id: $id})-[:ORDERED]->(o:Order)-[:CONTAINS]->(p:Product) RETURN count(*) AS times ORDER BY times DESC ``` Always use parameters (`$id`) in queries. This helps with both injection prevention and execution-plan caching. 💡 Use Cases It is the foundation for any domain that asks questions about connections: aggregating customer purchase history, recommendations, tracing paths in fraud detection, and walking org charts. It is highly readable and expresses complex traversals in fewer lines than SQL. ⚠️ Caveats ・Do not build queries by string concatenation; always parameterize. ・Watch for version differences. New features center on Cypher 25, and Cypher 5 is frozen. Verify compatibility. ・Because it is declarative, how you write a query can change the execution plan significantly. Check heavy queries with PROFILE/EXPLAIN. #Neo4j# #Cypher#
Show more