-
Constitutional AI: Harmlessness from AI Feedback
Paper • 2212.08073 • Published • 4 -
Specific versus General Principles for Constitutional AI
Paper • 2310.13798 • Published • 3 -
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Paper • 2501.18837 • Published • 10 -
Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
Paper • 2601.04603 • Published
Jash Shah
jash0803
AI & ML interests
Multi-Agents, Reinforcement Learning & Computer Vision
Recent Activity
upvoted a collection 3 days ago
Papers: RLAIF updated a collection 3 days ago
Papers: RLAIF updated a collection 3 days ago
Papers: RLAIF