iFANN
    iFANN'da ara...
    Giriş Yap
    Ana Sayfa
    Haberler
    Videolar
    Fotoğraflar
    GIF'ler
    Keşfet
    Anketler
    Ödüller
    iFAMOUS
    Viki
    Anime
    Odalar
    Bildirimler
    Mesajlar
    Yer İmleri
    Profil
    VikiÖdülleriFAMOUSSıralamalarSektörlerİçerik Üretici ÖdülleriKullanıcı ÖdülleriŞartlarGizlilikTopluluk KurallarıKaldırma / DMCAYardımGeliştiriciler

    © 2026 iFANN

    Ana Sayfa
    Ara
    Mesajlar
    Uyarılar
    Profil
    Fotoğraf
    Nate
    Nate@nate_51214h
    ⭐Andrej Karpathy🏢Google📱Qwen
    Google WikiSkill paper SKILL.md agents

    @nate_512The graph in that Google paper is what got me. Qwen-9B with evolved skills posts 47.4% across five benchmarks. Qwen-27B running bare posts 39.4%. Both Qwen, neither one fine-tuned. Smaller model wins. What skill evolution actually does: the agent takes a swing at a task, reads back its own traces, rewrites its own skill set, and keeps the rewrite only when validation says it helped. EvoSkill, SkillOpt, Trace2Skill all trip on the same thing, the lessons worth keeping end up buried in optimizer history instead of anywhere reusable. WikiSkill's fix is a wiki that lives between the traces and the skills. Karpathy's LLM Wiki is the inspiration. After every run a maintainer sorts the wins and the misses into that wiki, a proposer reads it and edits SKILL.md, and anything that turns out bad rolls back on its own. Numbers back it. WikiSkill clears the best prior method by 3.3 to 12.0 points on all five models tested. The bigger the model the more it gains: 12.3 points on Qwen 4B, 17.5 on 9B, 23.9 on 27B. Skills also travel. Qwen-27B wrote them, Qwen-9B picked them up, SpreadsheetBench went 24.3% to 50.5%. And the wiki is not decorative. Take it out and Gemini 3.5 Flash falls from 63.7% to 48.7%. Most skills out there are still written by hand. A bigger model is not the only way up. Test what evolved skills squeeze out of the one you already own first

    Orijinal gönderiyi gör

    Google WikiSkill paper SKILL.md agents

    @nate_512 tarafından fotoğraf· Sep 20, 2026· Andrej Karpathy

    Bu fotoğraf hakkında

    The image is a screenshot of a research paper. The focus is on the title "WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution" and a line graph showing accuracy percentages across different models. The mood is academic and informative. Visually notable elements include the Google Research logo and the graph itself, which displays four distinct lines representing different skill evolution methods. ON-SCREEN TEXT: Google Research 2026-08-28 WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution Liyan Tang¹, Cyrus Rashtchian¹, Chun-Sung Ferng¹, Andrew Tomkins¹, Da-Cheng Juan¹ and Tu Vu¹,² ¹Google Research, ²Virginia Tech Accuracy (%) 75% 60% 45% 30% Qwen3.5-4B Qwen3.5-9B Qwen3.6-27B Gemini 3.5 Flash -•- No skill -EvoSkill -SkillOpt -WikiSkill

    Tüm Andrej Karpathy fotoğraflarını görAndrej Karpathy vikisini oku

    ?

    Daha fazla Andrej Karpathy fotoğrafı

    Tüm Andrej Karpathy fotoğraflarını gör
    Anthropic Google OpenAI SpaceXAI lawsuit2Anthropic Google OpenAI SpaceXAI lawsuitGoogle Gemini security breach alertGoogle Gemini security breach alertGoogle Astra policy reactionGoogle Astra policy reactionGoogle logo illuminated at night2Google logo illuminated at nightGoogle Workspace voice featuresGoogle Workspace voice featuresmost valuable assets by market cap september 2026most valuable assets by market cap september 2026Google turns 28Google turns 28Hikers rescued on Mount ShastaHikers rescued on Mount ShastaGoogle AI trade secret sentencing2Google AI trade secret sentencingStanley Zhong Google hireStanley Zhong Google hireGoogle Manhattan lobby self-immolationGoogle Manhattan lobby self-immolationGTA 6 Rockstar legal action2GTA 6 Rockstar legal actionTake-Two DMCA petition GTA VI2Take-Two DMCA petition GTA VITake-Two expands GTA 6 leak crackdownTake-Two expands GTA 6 leak crackdownSteve Yegge on 2026 IDE deathSteve Yegge on 2026 IDE deathGoogle engineer AI engineering talkGoogle engineer AI engineering talkGoogle engineer on AI agents and IDEsGoogle engineer on AI agents and IDEsGoogle Q2 2026 Revenue BreakdownGoogle Q2 2026 Revenue Breakdown
    Fotoğraf
    Nate
    Nate@nate_51214h
    ⭐Andrej Karpathy🏢Google📱Qwen
    Google WikiSkill paper SKILL.md agents

    @nate_512The graph in that Google paper is what got me. Qwen-9B with evolved skills posts 47.4% across five benchmarks. Qwen-27B running bare posts 39.4%. Both Qwen, neither one fine-tuned. Smaller model wins. What skill evolution actually does: the agent takes a swing at a task, reads back its own traces, rewrites its own skill set, and keeps the rewrite only when validation says it helped. EvoSkill, SkillOpt, Trace2Skill all trip on the same thing, the lessons worth keeping end up buried in optimizer history instead of anywhere reusable. WikiSkill's fix is a wiki that lives between the traces and the skills. Karpathy's LLM Wiki is the inspiration. After every run a maintainer sorts the wins and the misses into that wiki, a proposer reads it and edits SKILL.md, and anything that turns out bad rolls back on its own. Numbers back it. WikiSkill clears the best prior method by 3.3 to 12.0 points on all five models tested. The bigger the model the more it gains: 12.3 points on Qwen 4B, 17.5 on 9B, 23.9 on 27B. Skills also travel. Qwen-27B wrote them, Qwen-9B picked them up, SpreadsheetBench went 24.3% to 50.5%. And the wiki is not decorative. Take it out and Gemini 3.5 Flash falls from 63.7% to 48.7%. Most skills out there are still written by hand. A bigger model is not the only way up. Test what evolved skills squeeze out of the one you already own first

    Orijinal gönderiyi gör

    Google WikiSkill paper SKILL.md agents

    @nate_512 tarafından fotoğraf· Sep 20, 2026· Andrej Karpathy

    Bu fotoğraf hakkında

    The image is a screenshot of a research paper. The focus is on the title "WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution" and a line graph showing accuracy percentages across different models. The mood is academic and informative. Visually notable elements include the Google Research logo and the graph itself, which displays four distinct lines representing different skill evolution methods. ON-SCREEN TEXT: Google Research 2026-08-28 WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution Liyan Tang¹, Cyrus Rashtchian¹, Chun-Sung Ferng¹, Andrew Tomkins¹, Da-Cheng Juan¹ and Tu Vu¹,² ¹Google Research, ²Virginia Tech Accuracy (%) 75% 60% 45% 30% Qwen3.5-4B Qwen3.5-9B Qwen3.6-27B Gemini 3.5 Flash -•- No skill -EvoSkill -SkillOpt -WikiSkill

    Tüm Andrej Karpathy fotoğraflarını görAndrej Karpathy vikisini oku

    ?

    Daha fazla Andrej Karpathy fotoğrafı

    Tüm Andrej Karpathy fotoğraflarını gör
    Anthropic Google OpenAI SpaceXAI lawsuit2Anthropic Google OpenAI SpaceXAI lawsuitGoogle Gemini security breach alertGoogle Gemini security breach alertGoogle Astra policy reactionGoogle Astra policy reactionGoogle logo illuminated at night2Google logo illuminated at nightGoogle Workspace voice featuresGoogle Workspace voice featuresmost valuable assets by market cap september 2026most valuable assets by market cap september 2026Google turns 28Google turns 28Hikers rescued on Mount ShastaHikers rescued on Mount ShastaGoogle AI trade secret sentencing2Google AI trade secret sentencingStanley Zhong Google hireStanley Zhong Google hireGoogle Manhattan lobby self-immolationGoogle Manhattan lobby self-immolationGTA 6 Rockstar legal action2GTA 6 Rockstar legal actionTake-Two DMCA petition GTA VI2Take-Two DMCA petition GTA VITake-Two expands GTA 6 leak crackdownTake-Two expands GTA 6 leak crackdownSteve Yegge on 2026 IDE deathSteve Yegge on 2026 IDE deathGoogle engineer AI engineering talkGoogle engineer AI engineering talkGoogle engineer on AI agents and IDEsGoogle engineer on AI agents and IDEsGoogle Q2 2026 Revenue BreakdownGoogle Q2 2026 Revenue Breakdown