fyi we had around ~10k new launches with the new pre-IPO feature release.
just imagine what would've happened if we didn't enforce ticker locks + client limits on new launches. we would've probably seen 100k launches within a single day.
this is where a platform becomes a product used by bots rather than real humans, who can't possibly filter through so many new assets. this ends up driving FOMO trades and chasing volume, which is not hard to inflate.
now also imagine what might be the effect of these launch limits constantly being tweaked with more demand coming in, esp once LONG will start scale very fast(hint concatenation of liquidity + flows into the best and most unique assets in the eco)
we are going to win so much that you'll get tired of winning.
next steps:
1. I want to apply filters that help users discover assets based on a variety of params (whale %, age, anti-fragility and some surprises). I think this is something we can innovate on as well. just better discovery.
2. we are actively monitoring any suspicious activity across all pairs. as I said, it's not hard to tweak our params. I don't see anything critical as of now, but we're already preparing some new params we can execute on with a click of a button, as soon as we feel like PNDs are becoming coordinated rather than organic market activity(remember AI nuked by 80%+ like 4 times)
3. As always, I think the best trading style on LONG is to find a team/project you think can cook and grow out of bad periods, help them out, and build a nice Longfolio of PVE assets. Keep in mind all of our top pairs were able to break ATHs or sustain value weeks post launch.
LONG.
Show more
# Weaviate Features and Practical Usage
🚀 Ever wished you could delete all the embedding-API plumbing from your app code? Weaviate's model provider integrations let you wire up vectorization, generation, and reranking just by writing it into your collection config.
📌 Title and Feature URL
Title: Model provider integrations
URL:
📝 Overview
Weaviate integrates with 20+ model providers including OpenAI, Cohere, Google, AWS, Azure OpenAI, Mistral, Anthropic, Hugging Face, and Ollama. You can plug them into automatic embedding at import, automatic embedding of query text, generation for RAG, and reranking of search results. The big win is that your application no longer needs code to call an embedding API and pass vectors in.
🔧 How It Works
Integrations fall into three roles:
- Vectorizer (embeddings): text or multimodal vectorization.
- Generative (LLM): text generation for RAG pipelines.
- Reranker: result refinement (offered by Cohere, Jina AI, NVIDIA, Voyage AI).
There are also two delivery forms. API-based providers (OpenAI, Google, Cohere, AWS Bedrock, etc.) call external services, while locally hosted options (Ollama, Hugging Face Transformers, Model2vec) run on your own infrastructure. API-based modules are enabled by default in v1.33+.
🛠 Practical Usage
- Specify an embedding provider with Configure.Vectors at collection creation, and Weaviate vectorizes automatically at both import and query time.
- Configure generation with Configure.Generative to run RAG over your search results.
- Configure a reranker with Configure.Reranker.
- Automatic vectorization targets text / text[] properties. Weaviate sorts property names alphabetically, concatenates them, optionally prepends the collection name, and sends the string to the model (you can also exclude properties per-field).
🎯 Use Cases
- Internal document search: auto-vectorize body text at import, and auto-vectorize the query with the same model so they stay consistent.
- Model swapping: change vendor or model by editing collection config only.
- Closed-network requirements: use a locally hosted option like Ollama to keep data in-house through embedding generation.
- RAG chat: combine retrieval and generation within the same configuration, minimizing external orchestration.
⚠️ Caveats
- API-based providers require API keys and incur usage charges.
- Rate limits follow each provider's policy; watch out during bulk imports.
- For versions before v1.27, the concatenated string is lowercased before being sent to the model.
- For versions before v1.33, set ENABLE_API_BASED_MODULES to use API-based modules.
#
Weaviate# #
Embeddings#
Show more
Tons of interesting things in the Kimi K3 tech report — here are five algorithm-side techniques that I think either I've never seen before or simply deserve more attention than they're getting.
1/ They open-sourced the model but kept the speculative decoding draft model, which could be a real serving advantage of their own.
Quick background for those less familiar — speculative decoding pairs a large target model with a small draft model. The draft cheaply proposes several tokens ahead, and the target verifies them all in one forward pass. The speedup is determined by the acceptance rate, i.e., how often the target agrees with the draft's proposals.
K3 is pre-trained with an MTP (multi-token prediction) layer, DeepSeek-V3 style — an extra layer on top of the backbone that predicts one token further into the future than the main next-token head. Structurally this layer is an exact copy of a regular backbone block, so you can think of it as a 94th layer that has been trained on the full pre-training corpus from day one, just with a shifted prediction target.
After post-training, they freeze the target and fine-tune this MTP layer into an EAGLE-3-style draft. The EAGLE family of methods makes the draft a single decoder layer that reads the target model's internal hidden features rather than only the generated token sequence — conditioning on the target's features is what lets a one-layer draft stay accurate.
The fine-tuning is then set up to match inference exactly. At inference, the draft proposes multiple tokens in a row, so from the second token onward it is building on its own unverified guesses rather than anything the target has confirmed. They replicate this condition during training by unrolling the draft for 7 steps — the first step uses the target model's features, and every step after that consumes the draft's own outputs from earlier steps.
Two more design choices worth knowing. The draft reads low/mid/high-level target features (outputs of the 1st, 4th, and final AttnRes blocks), concatenated and passed through a fusion matrix initialized as [0 0 I] — zero weights on the low and mid features, identity on the high-level one. At initialization the draft therefore sees exactly the high-level feature the MTP layer was pre-trained on, and it gradually learns to mix in the other two during fine-tuning. And instead of the usual KL surrogate, they directly minimize the negative log of the acceptance rate itself (the sum of min(p, q) over the vocabulary, where p and q are the target and draft distributions), since minimizing KL does not guarantee maximizing acceptance for a capacity-limited draft. Everything is trained under the same MXFP4/MXFP8 QAT as their serving stack.
The release itself is asymmetric. The full target weights are on HuggingFace, but the draft — and as far as I can tell, the MTP layer it was fine-tuned from — is not. That MTP layer was trained jointly with the backbone on the full pre-training corpus, which no external party has access to. So first-party serving stays faster and cheaper on the exact same open weights. This is the smartest business decision I've seen recently from open-weight model companies.
2/ Sync RL with partial rollouts.
Sync RL waits for every rollout in the batch to finish before updating, and since rollout lengths vary wildly, compute is wasted by waiting on all rollouts to complete. Async RL decouples actors from the learner, which is much more efficient, but actor weights go stale, and a long rollout can land several learner steps behind the current policy.
K3 runs a middle ground, where generation pauses as soon as a fraction of the trajectories completes, and optimization proceeds immediately, like async. However, unfinished trajectories get paused, enqueued, and resumed at the start of the next iteration under the freshly updated policy. In other words, a single 1M-token trajectory can literally be a relay across several different policy versions.
Essentially, they trade model staleness for data staleness — off-policy prefixes inside otherwise on-policy trajectories — and mitigate it with a per-token regularization that constrains each update to a localized neighborhood of the current policy. It's an intellectually pleasing trade.
What makes this viable at 1M context is the environment side, since pausing the model's rollout is easy; but pausing a live agentic sandbox mid-trajectory is not. Their microVM runtime checkpoints an environment in 133ms and resumes it in 49ms, and a paused sandbox consumes zero CPU and memory.
3/ Multi-teacher on-policy distillation as the merge step.
After RL, they have nine expert models — three domains (general / agents / coding) crossed with three reasoning-effort levels (low / high / max) — consolidated into one unified model through multi-teacher OPD. The idea itself is not new; the report cites the same lineage as Thinking Machines' OPD post, MiMo-V2-Flash, and DeepSeek-V4.
Three details stand out though.
First, this is distillation with zero compression. Teachers and student are the same 2.8T architecture, and OPD is purely the mechanism that folds nine RL policies into one model, not a way to shrink a big teacher into a small student.
Second, the OPD signal is implemented as a per-token RL reward — the clipped log-ratio between the teacher's and the student's probability of each generated token. Distillation is therefore not a separate pipeline; it is literally the same RL trainer running with a different reward. The student generates its own on-policy rollouts, the teacher scores every token along the way, and everything above carries over for free — partial rollouts, pausable sandboxes, the per-token regularization — which is what makes it feasible to distill even million-token agentic trajectories.
Third, a negative result. At each step the student samples one token from its own distribution and it is all the OPD reward looks at, requiring just a single number per step, the teacher's log-prob of that token. They experimented with finer-grained top-k objectives that match more of the teacher's distribution over candidate tokens at each step, and saw no advantage in either convergence speed or final performance. So they decided that no full logits were needed.
4/ The RL harness is randomized.
They represent an unified agent harness with shared tool interfaces, system prompts, context management strategies, skills, memories, subagents, and can instantiate Kimi Code, Claude Code, Codex, OpenClaw, Hermes, or entirely new harnesses from the same abstraction.
During RL, harness configurations are dynamically reshuffled across task groups so the model never overfits to any single tool schema or interaction protocol.
The implication is that harness generalization is a trained property, not an emergent one. If you have ever evaluated open models across different agent scaffolds and wondered why some transfer well and some fall apart, this is probably a big part of the answer.
It also fits Kimi's position as an open-weight company. A closed lab ships the model and the harness together and controls the whole stack; an open model gets dropped into whatever scaffold people already use — Claude Code, Codex, OpenClaw, some custom internal agent, so harness robustness is even more importnat for open-weights models. Interestingly, their own in-house coding bench even reports K3 scoring slightly higher under Claude Code than under their own Kimi Code.
5/ NoPE on every global attention layer.
All 24 Gated MLA layers in K3 use no positional encoding at all. Positional and recency information is carried entirely by the KDA layers' gating and decay (the backbone runs 3 KDA per 1 MLA), while the MLA layers do pure content-based global lookup.
The payoff shows up at context extension. K3 grows from 8K to 64K during pre-training and from 256K to 1M during cooldown with zero positional-encoding modification — no RoPE base retuning, no YaRN. Hybrid linear attention is usually pitched as the efficiency component of these architectures; here the linear layers are also doing the entire job of the position encoding.
Honestly, this only scratches the surface. The infra sections (MoonEP, quantile balancing, KDA-aware prefix caching) each deserve a post of their own. Full report is definitely worth the read.
Show more