Register and share your invite link to earn from video plays and referrals.

Search results for PlanningBench
PlanningBench community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including PlanningBench
Planning is where LLMs move from “saying” to “doing.” Tencent Hy, in collaboration with the Gaoling School of Artificial Intelligence at Renmin University of China, is excited to open-source PlanningBench - a scalable, verifiable framework for evaluating and training LLM planning capabilities. With PlanningBench, you get: ✅ 30+ real-world planning tasks ✅ Automated verification ✅ Evaluation and training support See how top-tier LLMs perform on PlanningBench 👇 Resources: arXiv: GitHub: HuggingFace: #PlanningBench# #TencentHunyuan# #OpenSource# 📷
Show more
continuous learning at 2 billion parameters on a single device with a custom personal-planning env this is a fun way to tinker with continuous learning problems. I've setup continuous calendar env. a small model plans a synthetic day, gets interrupted by a meeting, and has to repair the plan without reshuffling everything. the current setup: - OpenEnv for the environment. Costomised and preference based version of CalendarGym - explicit preferences and rule-based scoring instead of an LLM judge - SFT on corrected choices, mixing older examples into later training rounds - evaluate on the same held-out days after each round after 3 rounds: half as many simulated correction flags than a frozen model with access to the same memory pool, across 384 test days and 3 runs. the best result is SFT + replay, still can't get SDPO to work. a flag means another offered option scored better, not that a human had to intervene. still a toy, but now it’s a personal-planning bench i can just iterate on. you can run the environment locally and plug in your own policy: see below
Show more