Register and share your invite link to earn from video plays and referrals.

Search results for deeplearning
deeplearning community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including deeplearning
⚡ TL;DR: Can you instantly say how many times faster your model could run on your hardware? SOLAR derives the theoretical best-case runtime (Speed-of-Light) automatically and with validation, straight from PyTorch/JAX code. Title: SOLAR: AI-Powered Speed-of-Light Performance Analysis URL: 📌 Key points ・🤖 An LLM translates code into an IR, then a "generate-then-verify" loop checks it by numerically comparing outputs against the original ・🧮 A deterministic backend derives compute and memory traffic in closed form from just the einsum subscripts ・🎚 Multi-fidelity roofline bounds at three levels: unfused, fused, and cache-aware ・✅ 100% operator coverage on KernelBench's 270 problems with zero SOL violations (existing FLOP counters hit 75-84%) ・🚀 Headroom reaches 54.6x at L3, with fusion analysis surfacing another 7.8x of room ・🦾 All three robotics models on Jetson Thor were memory-bound; 500Hz control needs 19.7x current bandwidth ・🔄 Inverse roofline back-solves the minimum hardware specs needed to hit a latency target 🌍 Takeaway: Blending the flexibility of generative AI with the rigor of analytical math, it's a practical foundation for accelerating performance tuning and hardware selection without physical access. #DeepLearning# #Performance#
Show more
🧠 Biologically plausible learning without backpropagation fails on deep networks, and the real culprit is gradient rank collapse. Fixing it with geometric interventions lifts accuracy from 1.4% to 46.1%. Title: Overcoming Rank Collapse in Feedback Alignment URL: 📝 Overview Feedback alignment (FA) is a biologically plausible alternative to backpropagation. It avoids symmetric weight transport, using the same forward weights in the backward pass, and instead learns with fixed random feedback weights. This paper identifies why FA fails to scale to deep architectures, gradient rank collapse, and proposes how to overcome it. ❓ Challenges Solved The FA error has considerably lower rank than backpropagation and is confined to a lower-dimensional subspace. This rank collapse severely limits parameter-space exploration and is the root cause of weight alignment failing to progress during training. 💡 Methodology & Proposed Approach It proposes two complementary mechanisms to raise gradient dimensionality. ・The Muon optimizer, based on orthogonalization, sets every singular value to 1. It flattens the spectrum of momentum updates to emphasize previously neglected directions, rewriting the update geometry itself ・Batch normalization normalizes hidden-layer activations to promote activation orthogonality and maintain high-dimensional representations across layers ・Muon equalizes update geometry while batch norm preserves representational dimensionality; together they maintain richer learning directions throughout training, so alignment emerges even in deeper networks 🎯 Use Cases It is relevant to backprop-free learning, brain-inspired neuromorphic computing, and research on optimization trajectory dimensionality in general. 📊 Experimental Results ・It was evaluated on CIFAR-10, CIFAR-100, STL-10, and Tiny ImageNet with AlexNet and ResNet-18 ・On CIFAR-100 with ResNet-18, baseline FA reached just 1.4%, while FA plus batch norm hit 37.1% and FA plus Muon hit 25.3% ・FA plus Muon plus batch norm reached 46.1%, about a 9-point improvement over batch norm alone ・Similar gains were confirmed across all datasets and architectures #DeepLearning# #Neuroscience#
Show more
Harvard University just voted to limit the number of A grades given in undergraduate classes to about 20% of the class. I’m not in favor of this. It deeply runs counter to how I believe education should be. We should hold a high bar, but also work mightily to support the success of 100% of learners, rather than a fraction. Harvard’s administration took this step — over the objections of a large fraction of the student body — to counter grade inflation. Grade inflation is real: Many universities have been awarding A and B grades to ever larger fractions of students, and this has caused grade point averages (GPAs) to become less useful as signals of student skill. At the same time, we want students to succeed. The heart of the question is the role of educational institutions. Should our goal be: - To help students succeed? - To judge students? Both of these have value. But my focus when working in education is almost entirely helping students succeed. To me, it is clear that many people want to learn, to be empowered, to build skills that let them do new things! This is what we focus on at DeepLearningAI. This philosophy is also why my online courses (going back to my early online Stanford courses on Coursera) permitted an unlimited number of retries for graded assignments. I believe in letting — and even encouraging — someone to redo something until they succeed. This is as opposed to standing in judgement of the fact they didn’t get it right the first time. Further, I want homework assignments to be designed primarily to help people practice and learn, rather than to judge their skill level. This is why I prefer to create “Practice Problems” and “Practice Labs” — questions that, when you think through them, help you to gain practice and reinforce what you know. As opposed to “Assessment Problems” designed primarily to judge skill. But won’t Harvard’s move make GPAs more meaningful and help prospective employers identify strong candidates? Having hired a large number of people from Harvard and other institutions, I can say confidently that GPA is not an important signal. We have screening and interviewing processes that give far more accurate ways to figure out if someone is truly skilled. I do not need a wider spread in applicant GPA scores to figure out who's really good! To be clear, there is also value in assessment. Even though standardized testing is much hated, high-quality tests like the SAT, ACT, GRE, TOEFL, etc. provide objective measures of ability in a domain. I find that most people want to learn and succeed. There are also people who want rigorous assessment (for example, to apply for school admissions), but this is a lesser need, and is not my focus when building educational products. Harvard is often described as an “elite” educational institution. There are two ways to be elite: One option involves limiting enrollments, and then even among admitted students, cap the number of people that do well at 20%. I would rather pursue a different path: Set a high bar and teach elite, cutting-edge skills, but strive relentlessly to help everyone succeed. This way, eliteness is defined not by excluding people but by helping as many people as possible to be excellent. [Original text: The Batch newsletter]
Show more
0
194
2.2K
221
Forward to community
Yoshua Bengio, Turing Award winner and one of the founders of deep learning, on the question the field keeps postponing. He built the foundations of this technology. That he's now among those to ask where it leads tells you how open the question really is. #VoicesOn#
Show more
𝗣𝗮𝗶𝗱 𝗖𝗼𝘂𝗿𝘀𝗲 𝗙𝗥𝗘𝗘 (PART - 3) 1. Artificial Intelligence + Data Analyst 2. Machine Learning + Data Science 3. Cloud Computing + Web Development 4. Ethical Hacking + Hacking 5. Data Analytics + DSA 6. AWS Certified + IBM COURSE 7. Data Science + Deep Learning 8. BIG DATA + SQL COMPLETE COURSE 9. Python + OTHERS 10 MBA + HANDWRITTEN NOTES (72 Hours only ) Cost About - $500 To get: - 1. Follow (So I can DM you ) 2. Like & retweet 3. Reply " Send "
Show more
I'm deleting this soon because it's a legit cash-printing formula. 𝗣𝗮𝗶𝗱 𝗖𝗼𝘂𝗿𝘀𝗲 𝗙𝗥𝗘𝗘 (PART - 3) 1. Artificial Intelligence + Data Analyst 2. Machine Learning + Data Science 3. Cloud Computing + Web Development 4. Ethical Hacking + Hacking 5. Data Analytics + DSA 6. AWS Certified + IBM COURSE 7. Data Science + Deep Learning 8. BIG DATA + SQL COMPLETE COURSE 9. Python + OTHERS 10 MBA + HANDWRITTEN NOTES (72 Hours only ) Cost About - $500 To get: - 1. Follow (So I can DM you ) 2. Like & retweet 3. Reply " Send "
Show more
The field of artificial intelligence was officially born at the 1956 Dartmouth workshop, where John McCarthy coined the term “artificial intelligence.” Key founders include McCarthy, Marvin Minsky, Allen Newell, and Herbert Simon, who presented the first AI program, the Logic Theorist. Alan Turing laid the theoretical groundwork earlier with his 1950 paper and the Turing Test, asking if machines could think. So it’s more a group effort than one inventor. The original founders didn’t complete it at all. They set up the field and built early programs that solved math problems or played checkers, but the tech hit big limits. There were two “AI winters” where funding dried up because results didn’t match the hype. What we use today, like ChatGPT, comes from deep learning and neural networks that really took off around 2012 with AlexNet. That work was led by Geoffrey Hinton, Yann LeCun, and Yoshua Bengio, decades later. The Dartmouth group laid the vision, but modern AI is a completely different approach built on massive data and computing power they couldn’t dream of.
Show more
Everyone wants to build AI agents. Almost nobody has a reading list. They watch one YouTube video, skim a Twitter thread, then wonder why their agent falls apart the moment it needs to plan, use a tool, or remember anything. Here's the resource list I wish someone had handed me on day one. Every link, organized by what it actually teaches. Save this one. **VIDEOS** 1. LLM Introduction: 2. LLMs from Scratch: 3. Agentic AI Overview (Stanford): 4. Building and Evaluating Agents: 5. Building Effective Agents: 6. Building Agents with MCP: 7. Building an Agent from Scratch: 8. Philo Agents (playlist): **REPOS** 1. GenAI Agents: 2. Microsoft's AI Agents for Beginners: 3. Prompt Engineering Guide: 4. Hands-On Large Language Models: 5. GenAI Agents (alt link): 6. Made with ML: 7. Hands-On AI Engineering: 8. Awesome Generative AI Guide: 9. Designing Machine Learning Systems: 10. Machine Learning for Beginners (Microsoft): 11. LLM Course: **GUIDES** 1. Google's Agent Whitepaper: 2. Google's Agent Companion: 3. Building Effective Agents (Anthropic): 4. Claude Code Best Agentic Coding Practices: 5. OpenAI's Practical Guide to Building Agents: **BOOKS** 1. Understanding Deep Learning: 2. Building an LLM from Scratch: 3. The LLM Engineering Handbook: 4. AI Agents: The Definitive Guide — Nicole Koenigstein: 5. Building Applications with AI Agents — Michael Albada: 6. AI Agents with MCP — Kyle Stratis: 7. AI Engineering (O'Reilly): **PAPERS** 1. ReAct: 2. Generative Agents: 3. Toolformer: 4. Chain-of-Thought Prompting: 5. Tree of Thoughts: 6. Reflexion: 7. Retrieval-Augmented Generation Survey: **COURSES** 1. HuggingFace's Agent Course: 2. Build with Anthropic: I'm not telling you to go through all 39 links this weekend. Start with the Stanford overview, then Anthropic's "Building Effective Agents" guide, then ReAct. That sequence alone will teach you more about how agents actually work than most paid courses. Save this post. You'll want it back in three weeks when you're stuck debugging why your agent keeps calling the wrong tool. Repost if this saved you an afternoon of searching. P.S. Which one are you starting with?
Show more
You can boost generative AIs like Claude with Lean — not just during the spec phase, but also for security hardening and bug hunting after implementation. Here’s a cool result from recently. There was an attack vector in Plonky2-whir that even the strongest AI models kept missing no matter how hard they searched. We found it with Lean + Claude, so I turned it into a reusable skill.This skill is pretty beginner-friendly. You can just paste it into Claude or Codex and try it out even if you don’t know anything about Lean. I think the difference between using Lean to decide on specs or database design and then verify them, versus regular vibe coding, is kind of like building with stones or bricks versus using reinforced concrete. At first the speed feels similar, but eventually you reach a height where stacking stones just won’t cut it anymore, and the gap in scale and safety becomes really obvious. Lean is basically the reinforced concrete of software development.Until about 10 years ago, I honestly saw formal verification and logic-based stuff as something that got pushed to the sidelines by machine learning — like “not practical,” “doesn’t make money,” or “another failed expert systems thing.” Fuzzy logic was used in rice cookers and trains, and even in fighter jets later on, but pure logic systems like Prolog felt like they were no longer the main character after deep learning took over. So seeing Lean now working really well with machine learning and delivering these surprisingly strong results feels pretty special to me, since I’ve always liked the logic side of things.
Show more