Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

PaLM 2 on HumanEval coding benchmark (0 shot):

37.6% success

GPT-4:

67% success

Not even close, gpt4 miles ahead



GPT-4 is a fine-tuned model (likely first fine-tuned for code, then for chat on top of that like gpt-3.5-turbo was[0]), while PaLM2 as reported is a foundational model without any additional fine-tuning applied yet. I would expect its performance to improve on this if it were fine-tuned, though I don't have a great sense of what the cap would be.

[0] https://platform.openai.com/docs/model-index-for-researchers


They also write about Flan-PaLM2 which is instruction fine-tuned, but still some ways off GPT-4.


Humaneval needs careful consideration though.

In the GPT-4 technical report, they reported contamination of humaneval data in the training data.

They did measure against a "non-contaminated" training set but no idea if that can still be trusted.

https://cdn.openai.com/papers/gpt-4.pdf


There were a few reasoning benchmarks that I noticed think they omitted a direct comparison since they weren't as competitive compared to GPT-4, and instead opted to just show the benchmarks comparing itself to other versions of PaLM or other language models

HellaSwag: GPT-4: 95.3%, PaLM 2-L: 86.8%

MMLU: GPT-4: 86.4%, Flan-PaLM 2-L: 81.2%

ARC: GPT-4: 96.3%, PaLM 2-L: 89.7%

(from: GPT-4 paper: https://arxiv.org/pdf/2303.08774.pdf)




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: