Register and share your invite link to earn from video plays and referrals.

Search results for touken_hanamaru
touken_hanamaru community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including touken_hanamaru
"Token demand growth will need to outpace declining token prices to support continued growth in investment spending. Frontier models are currently a key source of demand for hyperscaler compute. However, the rise of competitive open-source models has contributed to a decline in average token prices. Measures of frontier token demand slowed in July..." - Goldman
Show more
Token prices new record lows
0
70
1.1K
144
Forward to community
Token costs are collapsing Lower prices and more sophisticated uses are fuelling adoption -DB Reid
Token creators have earned over $100,000,000 in stock tokens, ETH, USDG, and more on Pons. $1B next ➡️
Token costs new lows. The whistleblowing and resignations from veteran Anthropic employees who been with the company for 3-5 days is about to 10x.
Token volume on @vercel AI Gateway has averaged double-digit weekly growth for 8 straight weeks. Last week it accelerated to +24.8%. It’s almost inconceivable. Infinite demand of intelligence.
Token usage just went vertical, and the power grid is not ready for it (Save this). This chart shows that agentic token usage on OpenRouter rose from roughly 0.51 trillion tokens on February to about 7.3 trillion by Augustwhich is a fourteenfold increase in six months. Human usage grew only about 2.8 times during that period, reaching approximately 1.4 trillion tokens and agents now consume roughly five times more than humans. This matters because agents use roughly fifteen times more tokens per request, since a single employee action can trigger dozens of model calls, tool invocations, and retries before a person sees the result. This is Jevons paradox in practice because every improvement in inference efficiency lowers the cost per token and unlocks new behaviors rather than reducing total spending. The scale is already substantial, as OpenRouter's weekly token consumption grew roughly nine thousand times since early 2024 to more than 90 trillion, with agents accounting for approximately 71% of usage. Goldman Sachs estimates that agentic AI could increase 2030 token consumption by twenty four times. The honest counterargument is that roughly 85% of agentic tokens come from cached prompts billed at lower rates, so revenue is growing more slowly than raw token counts suggest. However, cached tokens still require servers, memory, networking and electricity, which means the physical demand remains real. The infrastructure math is the most important part because US data center power demand climbing from 31 gigawatts in 2025 to 66 gigawatts in 2027. That requires capacity additions of 36.3 gigawatts in 2027, compared with actual additions of only 8.5 gigawatts in 2025, which represents more than a fourfold acceleration. The announced pipeline points toward 135 gigawatts, but only about 60% of scheduled capacity is expected to arrive on time. The capital commitments are enormous, since combined spending from Amazon, Microsoft, Alphabet, Meta, Oracle, and SpaceX is projected to exceed $1.3 trillion by 2027, while total global AI data center investment could reach $7 trillion by 2030. This is why the bottleneck is shifting from chips to electricity, and some estimates project that 40% of AI data centers will face power shortages by 2027. The beneficiaries therefore extend beyond GPUs into memory, networking, power generation, transmission, cooling, and electrical equipment. If you want to see how Milk Road is positioning around this next wave of AI infrastructure spending, from power and cooling to memory and networking, check out the link below. We’re already making trades around the companies we think benefit most as agentic AI pushes demand even higher.
Show more
Token is the unit. Attention is the operation. The KV cache is the memory that lets attention reuse the past. Here is the map that connects all three. An LLM writes one token per pass through the model, from top to bottom. Pass 1 is prefill. The whole prompt goes in at once. At every layer, each position produces q, k, and v. Attention matches queries to keys and blends values. An MLP follows, and a new hidden state moves to the next layer. Each layer saves its K and V in the cache. At the bottom, the hidden state at the final position becomes scores over the vocabulary. Under greedy decoding, the highest-scoring one becomes the first new token. That token feeds back, and decode begins. Every pass after the first carries only the new position. At each layer: read the weights, compute the new q, k, and v, read the past K and V, run attention over the cached history plus the new pair, run the MLP, and append the new k and v to the cache. Earlier tokens never move through the layers again. Their cached K and V preserve what attention needs to reuse from them. That is what turns inference into a hardware problem. During prefill, one weight read from HBM can serve many positions in the prompt. During decode, the same read may serve only one new position. Add a growing cache that must be revisited on every pass, and generating a single token can require an enormous amount of data movement. Capacity, bandwidth, and locality often set the limit, not arithmetic alone. A long context window is therefore a memory budget before it is a product feature. Serving systems are also beginning to place prefill and decode on different hardware, because the two phases put very different demands on the machine. Keep the map. Much of modern AI infrastructure hangs from it. Quantization changes how weights and caches are stored and moved. GQA shrinks the cache. FlashAttention changes data movement. MoE changes which weights are read. Speculative decoding changes the loop itself. The same map extends to batching, latency, throughput, context length, memory hierarchies, interconnects, serving architecture, and the chips built to run it all. Different techniques. Different tradeoffs. The same machine underneath.
Show more
Token creators have now earned $34,000,000 on Pons Fees are earned in stock tokens, ETH, USDG, or other supported tokens.
Token Prices Are Falling Faster Than Demand, Goldman Warns On The Broad AI De-Rating