← Back to explorer

ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators

Type
other
Venue
arXiv / Stanford University / Google Brain

Summary

Replaced token detection: a small MLM generator fills 15% masks; the discriminator classifies every token as original vs replaced (MLE generator, not adversarial). Learns from all positions. ELECTRA-Small (14M, 4 days on 1 V100) GLUE 79.9 vs BERT-Small 75.1 and GPT 78.8. ELECTRA-Base 85.1 vs BERT-Base 82.2. ELECTRA-400K Large (335M, ~1/4 RoBERTa compute) GLUE 89.0 vs RoBERTa-500K 88.9; ELECTRA-1.75M 89.5 and SQuAD 2.0 test 88.7/91.4. ICLR 2020. Code https://github.com/google-research/electra.

Keywords

electra · replaced-token-detection · bert · glue · squad · iclr · google · stanford · pretraining

Topics

pretraining, replaced token detection, BERT

Research notes

  • Primary: arxiv abs (cs.CL). License: arXiv.org perpetual non-exclusive on HTML at check. Clark/Manning Stanford (CIFAR Fellow); Luong/Le Google Brain. Correspondence kevclark@cs.stanford.edu. Code on HTML https://github.com/google-research/electra (2,367 stars at check; Apache-2.0; archived). HF paper page 0 upvotes; 43 unofficial linked models not copied into hf_* fields. Discord posted the ICLR 2020 OpenReview PDF (id=r1xMH1BtvB); cataloged from the matching arXiv abs. Trains on public Wikipedia/BooksCorpus/ClueWeb/CommonCrawl/Gigaword rather than a new hosted corpus, so no datasets_local row. License field left blank per catalog convention.