Categories: FAANG

TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining

This paper was accepted to the ACL 2025 main conference as an oral presentation.
This paper was accepted at the Scalable Continual Learning for Lifelong Foundation Models (SCLLFM) Workshop at NeurIPS 2024.
Large Language Models (LLMs) trained on historical web data inevitably become outdated. We investigate evaluation strategies and update methods for LLMs as new data becomes available. We introduce a web-scale dataset for time-continual pretraining of LLMs derived from 114 dumps of Common Crawl (CC) – orders of magnitude larger than previous continual language modeling benchmarks. We also…
AI Generated Robotic Content

Recent Posts

Introducing Claude Fable 5.1 on AWS

Today, we’re excited to announce the availability of Claude Fable 5.1 on Amazon Bedrock and…

22 hours ago

The Range Rover Electric: Specs, Price, Availability

After long delays, JLR’s biggest gamble with its Range Rover brand is here with huge…

23 hours ago

A new kind of AI that does its thinking cheaply without words

There may soon be a new kind of artificial intelligence in town, one that uses…

23 hours ago

Connect an AgentCore Runtime hosted MCP server to Amazon Quick

Model Context Protocol (MCP) servers allow foundation models to access external data and tools, supporting…

2 days ago

The Best Labor Day Mattress Deals on Beds We’ve Tried in Our Homes

It’s one of the best times of the year to buy a mattress, and our…

2 days ago

A “quantum bath” puts quantum entanglement on autopilot

Physicists have demonstrated a new way to entangle distant quantum bits without the constant measurements…

2 days ago