AI & ML interests

LLM, Agents, Evaluation, Evals, AI Quality, AI Security, Red-teaming

julien-cย 
posted an update 2 months ago
view post
Post
5639
who's working on an NVFP4 version of Kimi-K3?
  • 4 replies
ยท
davidberenstein1957ย 
posted an update 10 months ago
alexcombessieย 
updated a Space about 1 year ago
pierljย 
in giskardai/realharm about 1 year ago

RealHarm

#2 opened about 1 year ago by
PhunvVi
davidberenstein1957ย 
posted an update about 1 year ago
mattbitย 
in giskardai/realperformance about 1 year ago

Update README.md

#2 opened about 1 year ago by
davidberenstein1957

Update README.md

#2 opened about 1 year ago by
davidberenstein1957
davidberenstein1957ย 
posted an update about 1 year ago
view post
Post
430
๐Ÿšจ LLMs recognise bias but also reproduce harmful stereotypes: an analysis of bias in leading LLMs

I've written a new entry in our series on the Giskard, BPIFrance and Google Deepmind Phare benchmark(phare.giskard.ai).

This time it covers bias: https://huggingface.co/blog/davidberenstein1957/llms-recognise-bias-but-also-produce-stereotypes

Previous entry on hallucinations: https://huggingface.co/blog/davidberenstein1957/phare-analysis-of-hallucination-in-leading-llms
  • 1 reply
ยท