Register and share your invite link to earn from video plays and referrals.

Search results for 0530ユカリゾーン被害者の会
0530ユカリゾーン被害者の会 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including 0530ユカリゾーン被害者の会
# Decision Points in AI Agent Development # Temperature 🎯 The Hook Are you using the same temperature for every task in your agent? Temperature isn't just a "creativity knob." In agent systems, structured outputs, tool calls, and user-facing responses each need fundamentally different temperature settings. Using one value for everything is leaving performance on the table. 📋 Overview Temperature controls how "peaked" or "flat" the probability distribution is when an LLM selects its next token. Near 0, the highest-probability token wins almost every time, producing deterministic and stable output. Higher values flatten the distribution, allowing lower-probability tokens through, increasing diversity and creativity. In AI agent systems, the optimal temperature varies dramatically across contexts: generating structured output, assembling tool call arguments, and producing free-form text each call for different settings. Temperature should be treated as a dynamic variable that shifts with task type, not a single fixed constant. 🔍 Decision Points Temperature is primarily driven by task_variability -- how routine vs. creative the task is. The decision flow is straightforward 🧭 1. Structured output (JSON / function calling)? → 0.0-0.3 2. Accuracy-first (fact extraction, classification, summarization)? → 0.2-0.5 3. Dialogue, explanation, communication? → 0.5-0.7 4. Creative writing, brainstorming, candidate generation? → 0.7-1.0 Additionally, higher failure_cost pushes the temperature ceiling down, and high cost_sensitivity environments should account for retry cost increases from higher temperatures. 💡 Key Details Reference values by task type 📊 - Structured output (JSON / function calling): 0.0-0.3. Minimizing schema violations is the priority - Classification, extraction, data transformation: 0.0-0.2. Accuracy and reproducibility are paramount - Summarization, explanation, customer support: 0.5-0.7. Balance naturalness with accuracy - Creative writing, brainstorming, candidate generation: 0.7-1.0. Diversity is the source of value - Tool argument generation: 0.0-0.2. Precise function and argument names are non-negotiable - Planning and reasoning: 0.3-0.6. Some exploration helps, but maintain logical consistency Start structured output temperature at 0. If schema violations occur at 0, the problem is your prompt or schema -- never rely on higher temperature to "accidentally" produce correct output 🚫 ⚖️ Trade-offs Too low and conversations become robotic 🤖 The model returns identical answers to identical questions, giving users a "template response" impression. Best-of-N sampling also breaks down -- candidates become near-identical, costing N times more for essentially N=1 results. Too high and structured outputs start breaking 💥 JSON field names drift, types mismatch, hallucinations increase -- especially dangerous for proper nouns, numbers, and dates. Tool call instability and loss of reproducibility compound the problem. Monitor the retry cost impact of temperature changes. If schema violation rates exceed roughly 5%, consider lowering the temperature. 🛠️ Use Cases Vary temperature by pathway within a single agent 🔀 Planning steps at 0.3-0.5, tool argument generation at 0.0-0.2, user-facing responses at 0.5-0.7. When switching models, adjust temperature simultaneously for a natural fit. Using Best-of-N? You need to raise the temperature. Generating N=5 candidates at temperature 0 produces 5 near-identical outputs. For N>1, set temperature to 0.5-0.8 and let a Judge select the best from diverse candidates. Be careful combining temperature with top_p ⚠️ Adjusting both simultaneously creates multiplicative effects with unpredictable behavior. As a rule, tune one and leave the other at its default. #AIAgents# #SoftwareArchitecture#
Show more
Gemini 3.7 Flash — the CHEAP model — tripled its agent score (Terminal-bench 3.0: 5.4% → 14.9%) in THREE WEEKS not a new generation. a point release. 21 days. flash went 3.5 → 3.6 → 3.7 in ~3 months while the price got cut in half
Show more
Last wk, oil +9% & ylds +11-26 bps across 2/30 curve w/ S&P/Nas/R2K -0.3%/-0.7%/-2.4%. This wk, I am watching reaction to 1) oil/rates, 2) calls to slow down AI development & 3) Fed on 9/16. I remain on the cautious side till US mid-terms on 11/3. This weekend, the CEO of Anthropic called for a slowing of frontier model development over safety concerns. This follows comments along similar lines by the CEO of OpenAI to employees last week if other companies were willing to do the same thing. The fundamental issues I have with this is 1) foreign adversaries would welcome the US slowing down AI development, 2) I view this as an attempt to slow down open-weight model development which would help the market dominance of OpenAI and Anthropic which are currently in the lead and 3) I do not see other companies agreeing to anything that slows down progress catching up to these two market leaders. Having said that, I could see 3rd party evaluators to limit liability risk going forward and some sort of executive order from the White House. But I hope the longer-term result of these actions is broadly distributed personal AI capabilities for all individuals versus having it become concentrated in the hands of a few companies. Along this vein of AI competition, after releasing their paid API of Muse Spark 1.3 two weeks ago with open-weight versions coming later, $Meta launched their personal AI agent Muse last week with the stock gaining 5%. With 3.6 billion daily active users, a hit product could yield large results. Meta is increasingly showing other ways they can monetize their AI capex spend. This should help the stock to re-rate from a 17x CY27 PE to a multiple closer to peers trading in the low 20s. Meta Connect on September 23–24 is another potential catalyst given their leading frontier model Watermelon should be coming at the latest by October. On the front of broadly distributed AI capabilities, $AAPL stock gained 4% last week on their new product launch. The foldable Duo will provide a personalized AI agent in your pocket with a 50% larger screen than a Pro Max. I continue to see a big upgrade cycle next year. The change from a 4” screen to 5.5” screen with the iPhone 6 drove revenue growth from 7% in FY14 to 28% in FY15. The Android ecosystem has had a foldable Samsung phone since 2019. As for the Fed on Wednesday, I believe Warsh will raise by 25 bps and echo his hawkish statements from Jackson Hole on August 28th that “Price stability is not self-executing… 65 months of sustained, elevated inflation sits squarely with the Central Bank.” The ECB statement last week when they hiked might provide some hints: “For inflation excluding energy and food, the baseline foresees 2.5% in 2026, 2.6% in 2027 and 2.3% in 2028. Compared with June, the baseline projection for inflation in 2026 is unchanged, while it has been revised up for 2027 and 2028… The outlook remains highly uncertain, with risks to the upside for inflation and to the downside for economic growth.” In summary, my caution between now and the US mid-terms on 11/3 remains for reasons I have fleshed out in prior posts including: 1. Don’t fight the Fed: The market historically under-performs during a hiking cycle with the bond market discounting 2 raises by year-end and 3.5 raises by mid-June of 2027. 2. Seasonal headwinds: September is down -0.5% on average and up only 48% of the time since 1957. 3. Historical volatility: S&P drawdowns of 10% between 7/31 and 11/9 have occurred in the lead-up to mid-terms since 1990. 4. Regulatory friction: There is bipartisan pushback against datacenter expansion that could hurt the AI buildout in the near-term. 5. Geopolitical risk: Despite US efforts to de-escalate, I believe Iran drags out hostilities at least through the 11/3 US mid-terms, keeping oil prices elevated. 6. Macroeconomic pressure: Long-term government bond yields are hitting multi-decade highs for several countries, slowing down growth and providing a reasonable alternative to stocks. I believe in not fighting the Fed, the bond market or seasonality. I like the odds stacked in my favor which should improve at least seasonally following the mid-terms.
Show more
🧩 Qwen3.8-Flash-Next: A 6B-Active Preview of the Qwen4 Architecture On the same night Zhipu's GLM-5.3-Flash took over the timeline, @Alibaba_Qwen open-sourced Qwen3.8-Flash-Next — explicitly positioned as a preview of the Qwen4 architecture. Zhihu contributor Kitt在进化 argues that for people who actually run models locally, this is the more practical release of the night. The local-deployment groups he is in are, in his words, on fire. His take: the model is called 3.8, but the architecture is really Qwen4 in preview — and it borrows the best ideas from across the field. 1️⃣ Smaller, cheaper, and realistic for local deployment Flash-Next has 125B total parameters with only 6B active — far smaller than the 300B-class GLM-5.3-Flash. The API is priced at ¥1 input, ¥3 output, and ¥0.1 per cached million tokens, roughly matching DeepSeek-V4-Flash's off-peak rates. His comparison, per 1M tokens in RMB — input / output / cache hit: 🔹 Qwen3.8-Flash-Next: 1.0 / 3.0 / 0.1 🔹 GLM-5.3-Flash: 0.4 / 1.4 / 0.115 (limited-time promo rate) 🔹 DeepSeek-V4-Flash (off-peak): 1.5 / 4.5 / 0.05 2️⃣ A Qwen4 preview wearing a Qwen3.8 name The author points out this is the same play Zhipu made: GLM-5.3-Flash's architecture is also completely different from GLM-5.3. Shipping the new architecture as open weights early is deliberate pathfinding. Inference frameworks like vLLM and SGLang, plus the quantization toolchain, all need lead time to adapt before Qwen4 proper arrives. 3️⃣ An architecture that borrows from everyone The author reads Flash-Next as a synthesis of the field's best recent ideas. 🔹 Attention: Qwen's in-house QSA, built to balance throughput and speed on long context. 🔹 Knowledge: it absorbs DeepSeek's Engram line of work, packing prior knowledge into 51B of N-gram side parameters — high knowledge density at minimal compute cost. 🔹 Training: the Muon optimizer, popularized by Kimi, scheduled together with AdamW. 4️⃣ Half the active parameters, still overtaking At 6B active, Flash-Next is nearly half the size of DeepSeek-V4-Flash — yet Qwen's reported benchmarks show it overtaking that model on multiple coding and agent leaderboards. The author says he is not worried about real-world experience: the recent Qwen3.8-27B already proved itself in daily use, and this sits on the same foundation. 🔗 Full Reading: #Qwen# #Alibaba# #Qwen4# #OpenWeights# #LLM# #AIInference# #MoE#
Show more
In 2023, more than 5,000 attended the 30th Annual "OWN IT!" Baron Conference at New York's iconic Met Opera House. The 2024, 31st Annual "Building Legacy" Baron Conference on November 15th at The Met… features exceptional executives of SpaceX...MSCI...Arch Capital...and Red Rock Resorts...the Baron Capital team...incredible entertainment...and an ice cream cone served by former ice cream truck driver Ron... Oh yeah...good luck winning one of three awesome Tesla Y door prizes. All expenses, including door prizes...are paid by @BaronCapital...not our clients... Investors should consider the investment objectives, risks, and charges and expenses of the investment carefully before investing. The prospectus and summary prospectuses contain this and other information about the Funds. You may obtain them from the Funds’ distributor, Baron Capital, Inc., by calling 1-800-99-BARON or visiting Please read them carefully before investing. Portfolio holdings as a percentage of net assets as of June 30, 2024 for securities mentioned are as follows: Space Exploration Technologies Corporation - Baron Asset Fund (2.9%), Baron Fifth Avenue Growth Fund (0.9%), Baron Focused Growth Fund (10.3%), Baron Global Advantage Fund (6.1%), Baron Opportunity Fund (2.8%), Baron Partners Fund (13.2%*); Tesla, Inc. - Baron Fifth Avenue Growth Fund (3.2%), Baron Focused Growth Fund (8.6%), Baron Global Advantage Fund (3.3%), Baron Opportunity Fund (3.5%), Baron Partners Fund (28.9%*), Baron Technology Fund (2.5%); MSCI Inc. - Baron Asset Fund (0.5%), Baron Durable Advantage Fund (2.1%), Baron FinTech Fund (2.5%), Baron Focused Growth Fund (3.1%), Baron Growth Fund (9.8%), Baron Partners Fund (1.8%*); Arch Capital Group Ltd. - Baron Asset Fund (4.8%), Baron Durable Advantage Fund (2.1%), Baron FinTech Fund (3.0%), Baron Focused Growth Fund (6.4%), Baron Growth Fund (12.7%), Baron International Growth Fund (2.8%), Baron Partners Fund (9.5%*); Red Rock Resorts, Inc. - Baron Discovery Fund (1.6%), Baron Focused Growth Fund (3.9%), Baron Growth Fund (1.5%), Baron Partners Fund (1.5%*), Baron Real Estate Fund (1.7%), Baron Small Cap Fund (3.6%). *% of Long Positions. Portfolio holdings are subject to change. Current and future portfolio holdings are subject to risk. All expenses associated with this conference are paid by Baron Capital, Inc. No conference expenses are paid by Baron Funds. BAMCO, Inc. is an investment adviser registered with the U.S. Securities and Exchange Commission (SEC). Baron Capital, Inc. is a broker-dealer registered with the SEC and member of the Financial Industry Regulatory Authority, Inc. (FINRA).
Show more
just 0.3-0.5% of $neet will be worth millions soon
mda 0.5.3 is a thing of beauty
Full Time: Agama 5(0)- 4(0)Manica Diamonds Telone 5(0)- 3(0) Triangle United #ChibukuSuperCup# Follow the PSL WhatsApp Channel:
The cheapest MiCA-regulated crypto exchanges ↓ (bookmark for later) Revolut – 0% / 0.09% OKX – 0.08% / 0.1% Gate – 0.1% / 0.1% Bybit – 0.1% / 0.3% Bitvavo – 0.2% / 0.3% Blockchaincom – 0.2% / 0.4% Bitpanda – 0.3% / 0.3% Kraken – 0.3% / 0.4% Bitstamp – 0.3% / 0.4% Cryptocom – 0.3% / 0.5% Robinhood – 0.5% Bit2Me – 0.5% / 0.6% Coinbase – 0.6% / 1.2% eToro – 1% Swissquote – 1% / 1% Data from @DefiLlama.
Show more
Blog posts we used to have: “how I reduced latency across several hundred machines by 200 ms by removing two O(n)^2 hot paths” Blog posts we have now: “Here’s how to get the most out of /ultrathink in the latest release of Claudex 3.0.5”
Show more
0
30
1.3K
89
Forward to community