Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🏗️
Building on HF
Dipankar Sarkar
PRO
dipankarsarkar
3
339
49
Follow
domofon's profile picture
raxisvictory's profile picture
edithatogo's profile picture
31 followers
·
84 following
https://www.dipankar.cc
dipankarsarkar
dipankar
dipankarsarkar
AI & ML interests
Building the AI-native stack. Agents as infrastructure, safety as architecture, performance as plumbing. I publish the receipts: papers, datasets, demos.
Recent Activity
reacted
to
dejanseo
's
post
with 🔥
1 day ago
Can humans recognize AI generated text? Test: https://dejan.ai/ai-vs-human/ Results: https://dejan.ai/blog/ai-vs-human-test-results/
reacted
to
kanaria007
's
post
with 🔥
1 day ago
✅ Article highlight: Benchmark Publication Without Governance Inflation (art-60-274, v0.1) TL;DR: This article argues that a benchmark result is not a governance maturity claim. A score may be real, reproducible, and worth publishing—and still say nothing by itself about safety, deployability, assurance, institutional quality, or platform maturity. 274 treats benchmark publication as a discipline of comparability, disclosure, lifecycle limits, and anti-inflation. Read: https://huggingface.co/datasets/kanaria007/agi-structural-intelligence-protocols/blob/main/article/60-supplements/art-60-274-benchmark-publication-without-governance-inflation.md Why it matters: • prevents measured results from being inflated into safety or maturity claims • separates historical results from current comparability • makes scope, freshness, omissions, and unsupported readings visible • allows honest publication without requiring full platform assurance • treats narrower wording as trust discipline, not underselling What’s inside: • the publication triad: comparability, disclosure, and anti-inflation • bounded publication outcomes such as PUBLISHABLE, PUBLISHABLE_WITH_LIMITS, NOT_COMPARABLE, and NOT_PUBLISHABLE • benchmark publication profiles • comparability disclosure notes • public non-claims registers • inflation checklists for result-to-maturity, comparison-to-assurance, historical-to-current, and wording inflation Key idea: Do not say: “this system scored well, therefore it is mature, safe, or ready to deploy.” Say: “this result was observed under this benchmark and comparability frame, remains valid within these lifecycle and disclosure limits, and does not support these broader governance claims.” Better benchmark publication is not a louder score. It is a result that is harder to overread.
reacted
to
TravisMuhlestein
's
post
with 🔥
1 day ago
One of the more interesting questions in AI-assisted software development is whether increasingly detailed system prompts actually improve code quality. Johnathen Chilcher (Senior SRE at GoDaddy) explored that question through a large-scale benchmarking effort. What started as 1,458 Python benchmarks has grown into 212,000+ controlled evaluations across Python, Go, JavaScript, and C#, using Claude Haiku, Sonnet, and Opus. Several findings stood out: Across the original Python benchmarks, no prompt configuration consistently outperformed an empty prompt. Generic instructions such as "write clean code" or "follow best practices" often reduced performance rather than improving it. The information that consistently helped wasn't generic advice—it was project-specific context the model couldn't infer from training, including repository structure, build commands, coding conventions, and the current state of the codebase. Chain-of-thought prompting helped in Go and C#, but hurt performance in Python. Prompt tone mattered: encouraging language generally outperformed high-pressure or urgent wording. The effectiveness of prompting techniques varied across programming languages and models. One aspect I found particularly compelling is that the conclusions evolved as the benchmark grew. Earlier recommendations changed as additional data became available—a good reminder that empirical evaluation matters more than intuition. The broader implication is that we're moving beyond prompt engineering toward designing systems that automatically provide AI coding agents with the context they actually need to succeed. 📄 Full article: https://www.godaddy.com/resources/news/what-an-effective-ai-coding-prompt-looks-like I'd be interested to hear whether others have observed similar patterns across GPT, Gemini, Llama, DeepSeek, or other coding models.
View all activity
Organizations
dipankarsarkar
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
a dataset
2 days ago
brian-learns/cdx-cc-news
Viewer
•
Updated
2 days ago
•
1.46B
•
1.6k
•
5
liked
a model
2 days ago
TheStageAI/Qwen3.5-9B-GGUF
Text Generation
•
9B
•
Updated
13 days ago
•
730
•
2
liked
3 datasets
2 days ago
FrontisAI/NatureBench
Viewer
•
Updated
Jun 25
•
90
•
16.3k
•
11
FrontisAI/NatureBench-traces
Viewer
•
Updated
2 days ago
•
1.08k
•
1.43k
•
1
Abhisek987/deplab-dependency-compatibility
Viewer
•
Updated
2 days ago
•
24.9k
•
42
•
1
liked
a model
2 days ago
Vulcora/gemma-2-2b-it-codealpaca-alignmenttax
Text Generation
•
Updated
6 days ago
•
15
•
1
liked
4 models
3 days ago
deepreinforce-ai/Ornith-1.0-9B
Text Generation
•
1.47M
•
Updated
Jun 25
•
2.33M
•
•
502
cesun/SODA-Agent-Safety-Judge
Text Generation
•
4B
•
Updated
Jun 18
•
68
•
2
unsloth/Kimi-K3-GGUF
Image-Text-to-Text
•
2.8T
•
Updated
5 days ago
•
128k
•
278
dejanseo/gemotions
Updated
May 17
•
4
liked
a dataset
3 days ago
SoulInPsyAbstract/sipa-os-governance
Updated
about 2 hours ago
•
150
•
1
liked
a Space
3 days ago
Running
9
mm-ctx
🗂
9
mm CLI in your browser — fast, multimodal context for agents
liked
6 datasets
4 days ago
ArchCoder/llm-cold-start-benchmark
Viewer
•
Updated
9 days ago
•
141
•
64
•
1
ejcgan/hint-faithfulness-transcripts
Preview
•
Updated
4 days ago
•
63
•
1
eth-easl/swissai-serving-trace
Preview
•
Updated
5 days ago
•
247
•
3
Alibaba-NLP/SecRespond
Updated
4 days ago
•
90
•
2
rogue-security/coding-agent-security-benchmark
Viewer
•
Updated
6 days ago
•
332
•
76
•
1
sergiopaniego/pelican-svg-drawings
Viewer
•
Updated
5 days ago
•
169
•
120
•
1
liked
a dataset
5 days ago
Glint-Research/GCI_Bench
Viewer
•
Updated
3 days ago
•
5k
•
206
•
6
liked
a model
5 days ago
UniversalComputingResearch/Limen0.2B
Text Generation
•
0.2B
•
Updated
10 days ago
•
209
•
9
Load more