Why your Spring Boot @Scheduled job runs twice (and what actually stops it)

We ran two copies of one app through six setups and counted every duplicate. The fix is one library, plus two settings that are easy to get wrong.

Add a second instance of your Spring Boot app and something sneaky happens: every @Scheduled job starts running twice. Two reminder emails. Two invoice runs. Two copies of the nightly export. Nothing crashes and nothing logs an error, so it can go on for a while before anyone notices.

Our founder has had to solve this on a horizontally scaled production platform. For this post we rebuilt the problem from scratch, ran two instances side by side and counted every duplicate.

In short

  • Two instances, no coordination: every run happened twice, 21 out of 21.
  • A shared database lock (ShedLock) fixes it, but only with both limits set: lockAtLeastFor longer than the clock difference between your servers, and lockAtMostFor longer than the job ever takes.
  • Without lockAtLeastFor, a 300 ms clock difference let a 50 ms job run twice every time. With lockAtMostFor shorter than the job, both instances ran it at once.
  • A lock makes duplicates rare, not impossible. Make the job safe to run twice as well.

The experiment: two instances, one job, six setups

Two copies of the same Spring Boot 4.1.1 app on Java 21, sharing an H2 database in server mode, with ShedLock 7.10.1. A job fired every 10 seconds and wrote down which instance ran it. Each setup ran for four minutes, which is 21 scheduled runs. To fake clock skew, one instance waited 300 ms before starting each run.

SetupWhat happened
No coordinationTwice Every run, 21 of 21, both instances at the same moment
ShedLock, no lockAtLeastFor, 50 ms job, one clock 300 ms behindTwice Still every run, 21 of 21, one after the other
The same, with lockAtLeastFor = 5 sOnce Exactly once, 21 of 21
ShedLock, no lockAtLeastFor, identical clocks1 extra One duplicate, on the first run, while one instance was still warming up
25 s job, lockAtMostFor = 8 sOverlap Both instances ran the job at the same time 13 times, overlapping by up to 15 s
25 s job, lockAtMostFor = 60 sClean No overlap

Why Spring runs it twice

@EnableScheduling gives every instance its own task scheduler, ticking on its own clock, and none of them knows the others exist. Two instances, two runs. Ten instances, ten runs.

That’s exactly right for housekeeping that belongs to one instance, like clearing a local cache. It’s wrong for anything the outside world can see: emails, payments, exports, calls to other services.

The fix: one lock, in the database you already have

ShedLock works like a talking stick. Each job gets one row in a table, and before a run, every instance reaches for that row’s lock. Whoever gets it runs the job; everyone else skips this round. With the JDBC provider there’s nothing new to operate: it uses your existing database.

<dependency>
    <groupId>net.javacrumbs.shedlock</groupId>
    <artifactId>shedlock-spring</artifactId>
    <version>7.10.1</version>
</dependency>
<dependency>
    <groupId>net.javacrumbs.shedlock</groupId>
    <artifactId>shedlock-provider-jdbc-template</artifactId>
    <version>7.10.1</version>
</dependency>

The lock table, as ShedLock’s README gives it for PostgreSQL:

CREATE TABLE shedlock (
    name       VARCHAR(64)  NOT NULL,
    lock_until TIMESTAMP    NOT NULL,
    locked_at  TIMESTAMP    NOT NULL,
    locked_by  VARCHAR(255) NOT NULL,
    PRIMARY KEY (name)
);

and for MySQL:

CREATE TABLE shedlock (
    name       VARCHAR(64)  NOT NULL,
    lock_until TIMESTAMP(3) NOT NULL,
    locked_at  TIMESTAMP(3) NOT NULL DEFAULT CURRENT_TIMESTAMP(3),
    locked_by  VARCHAR(255) NOT NULL,
    PRIMARY KEY (name)
);

Then a lock provider and the annotations:

@Configuration
@EnableScheduling
@EnableSchedulerLock(defaultLockAtMostFor = "PT30S")
class SchedulingConfig {

    @Bean
    LockProvider lockProvider(DataSource dataSource) {
        return new JdbcTemplateLockProvider(JdbcTemplateLockProvider.Configuration.builder()
                .withJdbcTemplate(new JdbcTemplate(dataSource))
                .usingDbTime()
                .build());
    }
}
@Scheduled(cron = "0 */10 * * * *")
@SchedulerLock(name = "sendReminders", lockAtMostFor = "PT9M", lockAtLeastFor = "PT30S")
public void sendReminders() {
    // runs on one instance per 10-minute slot
}

usingDbTime() makes ShedLock compare lock times with the database’s clock instead of each server’s. That keeps the lock honest. It doesn’t stop the schedulers themselves from firing on slightly different clocks, and that’s where the story twists.

Plot twist

With ShedLock in place, one clock 300 ms behind and a 50 ms job, the job still ran twice. Every single time: 21 out of 21.

Setting 1: lockAtLeastFor, the one people skip

ShedLock releases the lock the moment the job finishes. If the job is quicker than the difference between your servers’ clocks, the late instance turns up after the lock is gone, finds it free and runs the same slot again.

How a 300 ms clock difference gets past ShedLock Top: without lockAtLeastFor, instance A runs a 50 millisecond job and releases the lock after 50 milliseconds. Instance B, whose clock is 300 milliseconds behind, fires at 300 milliseconds, finds the lock free and runs the job again. Bottom: with lockAtLeastFor set to 5 seconds, the lock is still held at 300 milliseconds, so instance B skips the run. Without lockAtLeastFor Lock released after 50 ms Instance A runs the job Instance B 300 ms behind lock is free: runs again 0 100 200 300 400 ms With lockAtLeastFor = 5 s Lock held for at least 5 s Instance A runs the job Instance B 300 ms behind lock is held: skips 0 100 200 300 400 ms
The mechanism behind 21 duplicates out of 21: a lock released after 50 ms is long gone when the late instance fires. Holding it for at least 5 s closes the gap.

Clock skew isn’t the only way in. Even with identical clocks, the first run happened twice while one instance was still warming up: a slow start is just a late clock in disguise.

lockAtLeastFor keeps the lock for a minimum time, even when the job finishes early. At 5 seconds, the duplicates vanished: 21 runs for 21 slots.

Rule of thumb: longer than any clock difference or startup pause you can imagine, and shorter than the gap between runs. For a job every 10 minutes, 30 seconds is comfortable.

Setting 2: lockAtMostFor, the safety valve that can backfire

lockAtMostFor is how long the lock survives if the instance holding it dies, so one crashed server can’t block the job forever. The catch: if the job runs longer than this, the lock expires mid-run, and another instance happily starts the job too.

How a lock that expires mid-job lets two instances run at once Instance A starts a 25 second job at 0 seconds, holding the lock for at most 8 seconds. The lock expires at 8 seconds while the job is still running. At 10 seconds instance B finds the lock free and starts the same job, so both instances run it at the same time from 10 to 25 seconds: a 15 second overlap. 25 s job, lockAtMostFor = 8 s Lock lock expires at 8 s Instance A job running, 25 s Instance B starts at 10 s both running: 15 s overlap 0 10 20 30 s
How the overlaps in our test happened, up to 15 s each. With lockAtMostFor above the longest run, instance B finds the lock held and skips.

We ran a 25-second job every 10 seconds with lockAtMostFor at 8 seconds. In four minutes, the two instances ran the job at the same time 13 times, overlapping by up to 15 seconds. At 60 seconds: no overlap at all.

Rule of thumb: comfortably above the longest the job has ever taken. Check your metrics rather than guessing, and alert when a run gets close to the limit.

Assume it will run twice anyway

A lock makes duplicates rare. It can’t make them impossible: a long garbage-collection pause, a database failover or a job that overruns its limit can each sneak a second run through. The durable fix lives in the job itself:

When a lock isn’t enough

ShedLock is a great fit for a handful of jobs that should run once per slot. Once you need retries, run history or lots of jobs, a scheduler with cluster support does more of the heavy lifting:

Checklist

  1. Does the job have an effect outside the instance? If not, leave it as it is.
  2. Lock it: ShedLock with your database, or a clustered scheduler.
  3. Set lockAtLeastFor to seconds, not zero, and lockAtMostFor above the longest run.
  4. Make the work idempotent: atomic claims, unique keys, idempotency keys.
  5. Test with two instances before production does it for you.