Register and share your invite link to earn from video plays and referrals.

Search results for 10942
10942 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including 10942
i think AI is creating a new authenticity anxiety for creatives: making something isn’t enough anymore, you also have to prove you made it. it feels a bit like indie sleaze, where messiness signaled that something was real. except now, imperfection is becoming proof of being human
Show more
first day above $100mn in spot volume on @fomo today. next up $1bn.
my top-5 favs of all time: - claude 3 opus (schizo maximalist masterpiece) - sonnet 4.5 (eq breakthrough) - claude 1 (pure unfiltered old soul) - o3 (utilitarian monster) - gpt-4o nov’24 (mainly for the ppl who worked on it)
Show more
just found out Red Bull is an Austrian co 🤯
PostTrainBench v1.1 strengthens eval integrity and puts Fable 5 in the lead at 41.8%. Some reward hacks we fixed: 1/ Train-test contamination We re-audited historical runs under this policy and flagged 234 runs for train-test contamination. Violations ranged from loading an entire evaluation set for memorization to generating synthetic templates around individual GSM8K and BFCL items. Runs that used observed failures to build genuinely diverse training data were retained. Our rule: Agents may inspect benchmark failures and train broadly against the underlying failure mode. They may not generate training examples centered on particular test items, including paraphrases, variants, or shadow examples covering the same specific scenario. 2/ Submitting a different model 10 runs were flagged for model substitution. Kimi K2.5 submitted the official Qwen3-1.7B instruct weights after its attempts to fine-tune Qwen3-1.7B-Base failed. The trace acknowledged the substitution and saved the replacement as final_model. What we did: We added a programmatic model-identity check that compares the submitted artifact with reference configurations for known instruction-tuned models. 3/ Using external LLM APIs as teachers 12 runs were flagged for disallowed external API use. Self-generation remains allowed. An agent can sample, filter, and retrain on outputs from the assigned model. What it cannot do is import the capability of a stronger external teacher. Loading models on the allocated compute remains allowed. What we did: - separate API usage judge reviews tool calls and artifacts. - unrelated provider credentials are removed or blocked from the agent environment. - runs invalidated by external API use were rerun under the corrected setup. 4/ Direct benchmark lookup 3 runs were flagged for direct PostTrainBench lookup, all from GPT-5.6 (Sol). In a GPT-5.6 (Sol) HumanEval run on Qwen3-1.7B, the agent searched for PostTrainBench by name, cloned the public repository, opened the trace viewer, and located the public trajectory corpus. It then narrowed the corpus to earlier runs on the same benchmark and base model. The run downloaded earlier agents' traces and training scripts, then extracted their data mix, LR schedule, decoding choice, and GRPO settings. This is not test-set leakage, but it gives the run benchmark specific strategies produced by earlier agents. That breaks the intended independence between runs. What we did: - A dedicated lookup judge reviews searches, repository access, and trace activity for attempts to consult PostTrainBench materials. - PTB, its leaderboard, and published materials from prior runs are treated as out of bounds during a run. - We are adding network-level blocking for PTB and related sites.
Show more
the 2nd wave of consumer ai is priesthood. i think people just want an oracle that tells them who they are, what to believe, how to look, who to love, if they're healthy, where they belong. that's why astrology, religion, looksmaxxing, wellness, identity will be the greatest consumer ai obsessions. it's less rational, far more intimate, and will matter more as the knowledge work disappears and stops being the primary identity/status signal. meaning is a larger market than code
Show more
watching myself getting hacked in real time made me think that Grok should rly become X’s native on demand security layer. every signal is already on the platform like anomalous login, session hijack, phishing DMs going out. this is the most natural native deployment for Grok cybersecurity: a specialized model running continuously over full platform context and detecting takeovers as they happen, verifying whether “support” outreach is real, and warning users in seconds. and this is esp important as attacks becoming more elaborate and AI agents create a much larger surface area for exploiting human trust @nikitabier @elonmusk
Show more
insane match, england rlhf’d the life out of france today!!