Leco Li's picture

Building on HF

Leco Li PRO

imnotkitty

·

kitty_im_404

AI & ML interests

None yet

Recent Activity

replied to their post 4 days ago

One week until DeepSeek V4 drops. Reports say it's a full multimodal model—native image, video, and text generation. And it puts serious pressure on OpenAI and Google. What are you hoping to see from this release?

reacted to SeaWolf-AI's post with 🚀 5 days ago

ALL Bench — Global AI Model Unified Leaderboard https://huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard If you've ever tried to compare GPT-5.2 and Claude Opus 4.6 side by side, you've probably hit the same wall: the official Hugging Face leaderboard only tracks open-source models, so the most widely used AI systems simply aren't there. ALL Bench fixes that by bringing closed-source models, open-weight models, and — uniquely — all four teams under South Korea's national sovereign AI program into a single leaderboard. Thirty-one frontier models, one consistent scoring scale. Scoring works differently here too. Most leaderboards skip benchmarks a model hasn't submitted, which lets models game their ranking by withholding results. ALL Bench treats every missing entry as zero and divides by ten, so there's no advantage in hiding your weak spots. The ten core benchmarks span reasoning (GPQA Diamond, AIME 2025, HLE, ARC-AGI-2), coding (SWE-bench Verified, LiveCodeBench), and instruction-following (IFEval, BFCL). The standout is FINAL Bench — the world's only benchmark measuring whether a model can catch and correct its own mistakes. It reached rank five in global dataset popularity on Hugging Face in February 2026 and has been covered by Seoul Shinmun, Asia Economy, IT Chosun, and Behind. Nine interactive charts let you explore everything from composite score rankings and a full heatmap to an open-vs-closed scatter plot. Operational metrics like context window, output speed, and pricing are included alongside benchmark scores. All data is sourced from Artificial Analysis Intelligence Index v4.0, arXiv technical reports, Chatbot Arena ELO ratings, and the Korean Ministry of Science and ICT's official evaluation results. Updates monthly.

replied to their post 5 days ago

The most popular OpenClaw tool has been released! 🥇 Claw for All: The ultimate all-rounder. Simplifies deployment for both devs & pros with a seamless web/mobile experience. 🥈 OpenClaw Launch: Speed is king. Deploy your apps in under 30 seconds with a single click. 🥉 ClawTeam: Skip the setup. Get pre-configured AI agent blueprints built specifically for OpenClaw. 4️⃣ vibeclaw: Local-first. Run OpenClaw in your browser sandbox in literally 1 second. 5️⃣ Tinkerclaw: The startup favorite. Zero-code platform to deploy, manage, and scale AI assistants. 6️⃣ ClawWrapper: The "last mile" tool. Simplifies the entire packaging and launch process. Which one are you adding to your stack? 🛠️ (Source: OpenClaw Directory)

View all activity

Organizations

upvoted an article 12 days ago

Article

What superpower does Kimi-K2.5 bring to the table?

24 days ago

•

4

upvoted an article 26 days ago

Article

Uncensor any LLM with abliteration

Jun 13, 2024

•

793

upvoted 3 collections about 1 month ago

HunyuanImage

4 items • Updated 2 days ago • 13

Qwen3

84 items • Updated Dec 31, 2025 • 1.71k

Deepseek Papers

Deepseek papers collection • 31 items • Updated about 11 hours ago • 331