Register and share your invite link to earn from video plays and referrals.

jiayu
@daidaijiayu
voice @openai prev @googlelabs
98 Following    746 Followers
Ah ha! Yes the diffusion Gemma
I ran some real, live evals on Jev vs DiffusionGemma-as-Jev (my patch for vLLM!) DiffusionGemma comes out as the winner, I think. Headlines: Is Jev faster than DiffusionGemma? No ❌ (API vs DGX Spark) Is Jev smarter than DiffusionGemma? No ❌ (they're roughly tied!)
Show more
Nice
We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. You can toggle this behavior in /config.
Show more
They are accented English voices! The best performance would be to also include that accent in your instruction prompt
We've added several new voices for the GPT-Live API - check them out in our new developer playground.
Haven’t got a chance to try Jev yet. One thing I do recently is to parallelly send astra a prompt to analyze 10K data. I am not sure how intelligent jev could be for these types of tasks. However a super old distributed system idea came up: how about if we send 5 similar but different prompts (maybe in different temp) parallel if jev is super fast and then get a higher accuracy. Maybe it is the harness way to scale the test time compute even jev like model cannot think . Or maybe we can train these parallel voting behavior into the model. Another thing is I still remembered in GDM they created the diffusion Gemini or Gemma while I was there. I am very interested to know if jev style + text diffusion can make what. I remember that model is super freaking fast. Fun thing
Show more
magic
you can use codex voice from your phone now, connecting to your computer from anywhere!! @axbehr and i got to star in this new codex ad, showing how much you can get done while working out. btw, this is powered by the new gpt-live-1
Show more
Created a GPT-Live-1 voice app template with @expo - Full-duplex audio + background calls on iOS - Choose from all supported voices - 6 Rive animations from Vercel AI Elements - Expo SDK 57 + native Expo UI controls Source below ↓
Show more
Impressive! Generating json is actually most of the agentic system doing day to day. I still remember in my time at Google, before json mode, we used an internal tool (dasm) that builds json schema to let Gemini generate JSON. And before that I tried to build a harness generating json field by field parallelly. Training the paradigm into the model is a great idea.
Show more
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
Show more
language learning! Nice use case
This is the closest @Joe_Speaking has felt to a real IELTS Speaking exam. Our current Gemini Live version is already smooth. With @OpenAIDevs’ gpt-live-1, the conversation flows even better, and the examiner keeps listening the whole time, even while it’s talking. You can just speak naturally! Raw demo, no edits. Hoping to ship by the end of Sep.
Show more
i kinda like this. enabling everyone has the frontier access on models especially the safety researchers should fit the mission better to "AGI benefits everyone".
Interesting one. When to speak and when to interrupt is fun.
We are releasing Duplex Cue today - a new evaluation dataset of in-turn adaptation behavior in full-duplex voice agents. Full-duplex evaluation often emphasizes whether an agent keeps speaking or stops during speech overlaps. That binary cannot express a third response humans use routinely: continuing to speak while incorporating what the listener just contributed. A full-duplex agent can keep talking through the cue, incorporate it without stopping, or hand over the turn. A simple stop-or-continue score cannot tell those behaviors apart. We introduce Duplex Cue, an evaluation of this in-turn adaptation behavior in full-duplex voice agents. Duplex Cue separates listener intent (backchannel, collaboration, or interruption) from speaker behavior: continuing unchanged, adapting within the turn, or yielding. Adaptation includes acknowledgment as well as content revision. In a single-model case study using 300 human-confirmed cues from unscripted English conversations, we compare recorded human responses with @nvidia 's PersonaPlex continuations generated while replaying the listener's audio. We retain 208 pairs with the ongoing speaker active at cue onset and a scorable response in each condition. For backchannels, PersonaPlex and recorded speakers show similar response patterns. For interruptions, PersonaPlex yields more often and adapts less often than recorded speakers. On the 66 collaboration pairs, recorded speakers adapt in 68.2% of cases, compared with 34.8% for PersonaPlex. The model otherwise continues unchanged (42.4%) or yields (22.7%). Natural human interaction is much more complicated than a simple stop-or-continue, and we are only scratching the surface of it! Findings, paper, and public dataset can be found at
Show more
The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot. We welcome a federal framework that sets consistent safety requirements for frontier AI. But we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence. Consistent rules to manage frontier risk so that we can maximize the benefits are a good idea (and we are excited by ideas like independent auditors). Years ago, companies like ours developed things like Responsible Scaling Policies and Preparedness Frameworks. Those were good for that moment, and focused primarily on the deployment of completed models, not what happens during their development process. Today's shift to focusing on safe development and evaluation will need new tools. For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability, in addition to the safety work we have long done in advance of model releases. We hope that other companies will learn from our approaches and propose their own; we think shared standards for misalignment, monitoring, and safety will lead to better outcomes. We look forward to collaborating with our colleagues across the industry to formulate the best version of these. When we talk about “pacing”, we do not mean “stopping”. Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs. Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness, or let capabilities get ahead of alignment and monitoring. Where we will need the help of our government is for international coordination. But first we should do what we can ourselves.
Show more
0
2.7K
14.2K
1.3K
Forward to community
Now powered by the GPT-Live SIP (telephony) API!
Another random idea: Is is possible to use AOSP and build a codex os. Basically natively integrate codex init. Enable native phone use and running it as the always on background assistant. Feels like similar to GrapheneOS but the goal is codex integration. Pixel phone is BL unlockable, probably it is totally doable.
Show more
been working on gpt-live for a while, and today we bring the SOTA to voice to everyone. you can build your own codex-voice in a super low price: 0.05/min. very proud to bring this to everyone, super natural, affordable and capable personal AGI
Show more
In my ~14 months at @OpenAI, one of the most surprising and genuinely wonderful things about it, has been the full leadership support for full-duplex models (gpt-live series), even when it seemed deeply improbable that they would work. And yet, if they did work, it was obvious how magical they could be. I am quite sure this is not an exception but a general rule for research projects at @OpenAI. There were so many moments when the problem felt impossibly hard. But one thing kept being true: it was never clear why it shouldn’t work. And almost every time we understood the problem a little better, it became a little easier to solve. I’m pretty sure there are very few environments in the world where a bet like this could have been made and sustained. So grateful to be part of this wonderful place.
Show more
0
52
1.3K
54
Forward to community
I installed opencode this weekend. I think muse spark and qwen 3.8 24B are two underrated model. Ox alpha is a bit overhyped, maybe too many people using it so it is slower than expected and Gemini 3.7 flash is really slow at agentic. OAI models are still the best though imo, it actually nails most of the actual work I did this weekend but just saying my impression on the free ones.
Show more
TBH, the fun thing is I used to roaming around OSS GitHub repo to find a good docker to do this and that in my homelab. Now I just build it in my flavor. Amazing
This weekend I tried to let my codex rebuilt my homelab, it spent 15% my weekly tokens. 1. Fixes all known CVE in my NAS + openWRT 2. Upgrade my alist to open list 3. Implemented the inference optimization for my local qwen 4. Implemented a mobile friendly open code UI 5. Created a really nice web based book reader to read things in my NAS And then @thsottiaux resets :)
Show more
This weekend I tried to let my codex rebuilt my homelab, it spent 15% my weekly tokens. 1. Fixes all known CVE in my NAS + openWRT 2. Upgrade my alist to open list 3. Implemented the inference optimization for my local qwen 4. Implemented a mobile friendly open code UI 5. Created a really nice web based book reader to read things in my NAS And then @thsottiaux resets :)
Show more
Finally got sometime to mess around with my homelab. It’s impressive on local model communities work. I can run qwen 3.8 at 110 TPS on a single 4090 now with llama.cpp. Truely goat and democratize intelligence.
Show more