Skip to Content
Key/Value Cache

Key/Value Cache

Every organization with the Key/Value Cache addon gets a private, Redis-compatible cache. Connect with any Redis client — redis-py , ioredis, go-redis or redis-cli — to memoise lookups, deduplicate work and share state between runs and parallel jobs. Your spiders get the connection URL with zero configuration.

It’s a cache, not a database: every key expires within 24 hours, and a key can be evicted earlier. Write your spider so that a miss simply means “do the work”. For data you need to keep, use Storage or the postgres add-on.

Why use it

A long-running spider walks categories and stores, and for every product it needs one extra request just to get the EAN. The same product shows up in many categories and stores, so that request repeats — and Scrapy’s duplicate filter drops the repeat, taking the item that was waiting for its EAN with it. The usual workaround is a giant dict in memory: it grows for the whole run, dies with the job, and isn’t shared with the jobs crawling in parallel.

With the cache, you look the EAN up first, fetch it only on a miss, and store it for the next product, the next run, and every parallel job. Typical uses:

  • Memoise enrichment lookups — EANs, geocoding, one field from a detail page.
  • Deduplicate work across runs and parallel jobs — “has anyone handled this URL today?”
  • Serve repeated downloads from a shared HTTP cache.
  • Counters and small shared state — INCR, hashes, sets.

Enabling it

  1. Enable Key/Value Cache under Billing → Addons in the dashboard.
  2. In Crawlers → Projects, turn on the cache add-on for each project that should use it.
  3. Add a Redis client to your image — for Python, redis>=5:
RUN pip install --no-cache-dir "scrapy>=2.12,<3" "redis>=5"

Every new job of that project then finds the connection in its environment.

Connecting

VariableWhat it is
INSIGHT_CACHE_URLredis://username:password@host:6379/0 — your organization’s cache
REDIS_URLThe same URL, under the name many libraries look for

If your project has a secret named REDIS_URL, your secret wins; INSIGHT_CACHE_URL always points at the cache.

import os import redis r = redis.Redis.from_url(os.environ["INSIGHT_CACHE_URL"], decode_responses=True) r.set("hello", "world", ex=3600) print(r.get("hello")) # world
  • Any client, default settings. RESP2 and RESP3 are negotiated automatically, so redis-py, ioredis, go-redis, node-redis and redis-cli all work out of the box.
  • Database 0 only. SELECT 0 works; any other number is rejected.
  • No TLS. The cache is reachable only from inside the platform (your jobs) and over the VPN for local development, so the URL is redis://, not rediss://.
  • One client per process. A client pools its connections; don’t create one per request. Your organization can hold 64 connections at once, across all jobs and machines.

Example: memoise a lookup

The EAN case from above. Look in the cache first; on a miss, fetch the EAN and store it for everyone else:

import os import redis import scrapy class ProductsSpider(scrapy.Spider): name = "products" def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) # One client per spider, shared by every callback. self.cache = redis.Redis.from_url( os.environ["INSIGHT_CACHE_URL"], decode_responses=True ) def parse_product(self, response): item = { "sku": response.css("[data-sku]::attr(data-sku)").get(), "name": response.css("h1::text").get(), } ean = self.cached_ean(item["sku"]) if ean: item["ean"] = ean yield item else: yield scrapy.Request( f"https://api.example.com/products/{item['sku']}/ean", callback=self.parse_ean, cb_kwargs={"item": item}, dont_filter=True, # never let the duplicate filter drop this item ) def parse_ean(self, response, item): item["ean"] = response.json()["ean"] try: self.cache.set(f"ean:{item['sku']}", item["ean"], ex=86400) except redis.RedisError: pass # budget exhausted or cache unreachable: the item is still complete yield item def cached_ean(self, sku): try: return self.cache.get(f"ean:{sku}") except redis.RedisError: return None # treat any cache problem as a miss def closed(self, reason): self.cache.close()

Inside the platform each cache call is a sub-millisecond network round trip, so the regular (blocking) client is fine in ordinary callbacks. dont_filter=True matters: the same SKU can be looked up twice before the first answer lands in the cache, and without it the duplicate filter would drop the second request — and its item.

With async def callbacks

If your project runs the asyncio reactor (TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor", the default since Scrapy 2.13), use the asyncio client and await it:

import os import redis.asyncio as aioredis import scrapy class ProductsSpider(scrapy.Spider): name = "products" def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) self.cache = aioredis.Redis.from_url( os.environ["INSIGHT_CACHE_URL"], decode_responses=True ) async def parse_product(self, response): item = {"sku": response.css("[data-sku]::attr(data-sku)").get()} ean = await self.cache.get(f"ean:{item['sku']}") # wrap in try/except as above if ean: item["ean"] = ean yield item else: yield scrapy.Request( f"https://api.example.com/products/{item['sku']}/ean", callback=self.parse_ean, cb_kwargs={"item": item}, dont_filter=True, ) async def parse_ean(self, response, item): item["ean"] = response.json()["ean"] await self.cache.set(f"ean:{item['sku']}", item["ean"], ex=86400) yield item

Example: deduplicate across runs and jobs

SET key 1 NX EX 86400 writes the key only if it doesn’t exist yet — and tells you whether you were first. It’s a single atomic command, so two jobs racing for the same URL can’t both win:

def parse_listing(self, response): for href in response.css("a.product::attr(href)").getall(): url = response.urljoin(href) # True only for the first job (or run) to claim this URL in the last 24h. if self.cache.set(f"claimed:{url}", 1, nx=True, ex=86400): yield scrapy.Request(url, callback=self.parse_product)

A claim means “someone took it”, not “it’s done”: if a job dies right after claiming, that URL is skipped until the key expires. To record completion instead, set the key after the item has been saved.

Locks

redis-py’s Lock helper (r.lock(...)) doesn’t work here. It acquires with SET … NX PX, but it releases with a Lua script (EVALSHA), which the cache refuses with ERR command 'evalsha' is not available on the InsightScrap cache — and the lock key then stays put until its timeout. Use the same claim pattern instead: SET … NX EX to take the lock, DEL to release it.

import uuid token = uuid.uuid4().hex # who holds the lock, handy when debugging if r.set("lock:daily-export", token, nx=True, ex=600): # held for 10 minutes at most try: run_export() finally: r.delete("lock:daily-export")

Give the lock an expiry longer than the work takes. If it expires first, another job can take the lock, and your delete would release theirs.

Example: a shared HTTP cache for Scrapy

Scrapy’s HTTP cache can keep responses anywhere. Drop this storage backend into your project and repeated requests are answered from the cache instead of the network — within the run, in later runs, and in parallel jobs.

# myproject/cache_storage.py import json import logging import os import redis from scrapy.exceptions import NotConfigured from scrapy.http import Headers from scrapy.responsetypes import responsetypes logger = logging.getLogger(__name__) MAX_TTL = 86400 # the cache never keeps a key longer than 24 hours MAX_BODY = 900 * 1024 # values are capped at 1 MiB; leave room for the rest class InsightCacheStorage: """Scrapy HTTP cache storage backed by the InsightScrap Key/Value Cache.""" def __init__(self, settings): self.url = os.environ.get("INSIGHT_CACHE_URL") or os.environ.get("REDIS_URL") if not self.url: raise NotConfigured("INSIGHT_CACHE_URL is not set") expiration = settings.getint("HTTPCACHE_EXPIRATION_SECS") self.ttl = min(expiration, MAX_TTL) if expiration > 0 else MAX_TTL def open_spider(self, spider): # decode_responses=False: bodies are stored and returned as raw bytes. self.r = redis.Redis.from_url( self.url, decode_responses=False, socket_timeout=2, socket_connect_timeout=2 ) self.fingerprinter = spider.crawler.request_fingerprinter def close_spider(self, spider): self.r.close() def _key(self, spider, request): return f"httpcache:{spider.name}:{self.fingerprinter.fingerprint(request).hex()}" def retrieve_response(self, spider, request): try: data = self.r.hgetall(self._key(spider, request)) if not data: return None url = data[b"url"].decode() status = int(data[b"status"]) headers = Headers({ k.encode("latin-1"): [v.encode("latin-1") for v in values] for k, values in json.loads(data[b"headers"]).items() }) body = data[b"body"] except (redis.RedisError, KeyError, ValueError) as e: logger.debug("httpcache read failed, treating as a miss: %s", e) return None respcls = responsetypes.from_args(headers=headers, url=url, body=body) return respcls(url=url, headers=headers, status=status, body=body) def store_response(self, spider, request, response): if len(response.body) > MAX_BODY: return # latin-1 round-trips any header byte; repeated headers (Set-Cookie) are kept. headers = { k.decode("latin-1"): [v.decode("latin-1") for v in values] for k, values in response.headers.items() } key = self._key(spider, request) try: pipe = self.r.pipeline() # MULTI ... EXEC pipe.delete(key) # a re-store starts a fresh lifetime pipe.hset(key, mapping={ "status": response.status, "url": response.url, "headers": json.dumps(headers), "body": response.body, }) pipe.expire(key, self.ttl) pipe.execute() except redis.RedisError as e: logger.debug("httpcache write failed, skipping: %s", e)

Turn it on in settings.py:

# settings.py HTTPCACHE_ENABLED = True HTTPCACHE_STORAGE = "myproject.cache_storage.InsightCacheStorage" HTTPCACHE_EXPIRATION_SECS = 86400 # anything longer is capped at 24h anyway

Then yield repeated requests with dont_filter=True. The duplicate filter runs before the HTTP cache, so a filtered request never gets the chance to be answered from it:

yield scrapy.Request(ean_url, callback=self.parse_ean, cb_kwargs={"item": item}, dont_filter=True)

Responses served from the cache carry "cached" in response.flags.

  • Every downloaded response is cached and counts toward your write budget. Keep big or one-off pages out with meta={"dont_cache": True}, and don’t cache blocks: HTTPCACHE_IGNORE_HTTP_CODES = [403, 429, 500, 502, 503, 504].
  • Bodies over ~900 KB are skipped (a single value can be at most 1 MiB).
  • If the cache is unreachable or your budget is used up, every lookup is a miss and the spider downloads as usual — the crawl never fails because of the cache.
  • Without INSIGHT_CACHE_URL (a local run with no .env), the backend switches itself off and Scrapy runs without an HTTP cache.

From your laptop

Reveal your credentials in the dashboard under Proxy Users → Cache (Redis-compatible). The card gives you the URL, host, port, username and password, a ready-to-paste .env block (INSIGHT_CACHE_URL=… and REDIS_URL=…), and a redis-cli command. Your machine must be connected to the VPN (WARP), as for the other local-development credentials.

export INSIGHT_CACHE_URL='redis://username:password@host:6379/0' # from the dashboard redis-cli -u "$INSIGHT_CACHE_URL" PING # PONG redis-cli -u "$INSIGHT_CACHE_URL" SET greeting hello EX 300 redis-cli -u "$INSIGHT_CACHE_URL" GET greeting # "hello" redis-cli -u "$INSIGHT_CACHE_URL" TTL greeting # (integer) 300 redis-cli -u "$INSIGHT_CACHE_URL" --scan --pattern 'ean:*' | head

Use redis-cli 6 or newer (older versions don’t send the username). Your laptop sees the same keyspace as your production jobs, so prefix the keys you write while experimenting (for example dev:).

Keys expire within 24 hours

The rule is simple: data lives at most 24 hours after it was written.

You doWhat happens
Write a key without a TTLIt gets 24 hours.
Write with a TTL over 24 hours (SET … EX/PX/EXAT/PXAT, SETEX, PSETEX)Silently clamped to 24 hours.
Create a key with HSET, SADD, RPUSH, INCR, APPEND, MSET, SETNX, GETSET or SET … KEEPTTLIt gets 24 hours if it has no TTL; an existing TTL is left alone.
EXPIRE, PEXPIRE, EXPIREAT, PEXPIREATCan only shorten the TTL (like the LT flag); anything over 24 hours is clamped first. A request to lengthen returns 0 and changes nothing.
EXPIRE … GT or PERSISTRejected.
> SET a 1 OK > TTL a (integer) 86400 > SET b 1 EX 604800 # asked for 7 days OK > TTL b (integer) 86400 # clamped to 24h > EXPIRE b 3600 (integer) 1 # shortened > EXPIRE b 7200 (integer) 0 # can't lengthen

What that means in practice:

  • Overwriting a key with SET starts a fresh lifetime (up to 24 hours).
  • Adding to a hash, set or list does not extend it. The key still expires 24 hours after it was created. For rolling data, overwrite it or use per-day keys (seen:2026-10-05).

Write budget

Each organization can write up to 1 GiB per rolling 24 hours — counted as the bytes of keys, fields and values you send in write commands, plus a small fixed overhead for what storing them really costs: about 100 bytes per key and 16 bytes per field, member or value. Reads and deletes don’t count against it.

In practice that overhead only matters for very small entries: a million SET product:123 <13-digit EAN> writes cost about 140 MB of budget, not the 24 MB of raw key and value bytes.

  • Over budget, write commands fail with an error starting with OOM (OOM cache write budget exceeded for this organization …). In redis-py that’s redis.exceptions.OutOfMemoryError, a kind of ResponseError.
  • Reads and deletes keep working, and the budget frees up as older hours roll out of the 24-hour window.
  • Deleting doesn’t give budget back. The budget counts bytes written in the window, not bytes stored, so DEL, FLUSHDB or overwriting a key frees none of it — only time does.
  • Because nothing lives longer than 24 hours, the budget is also the most data your organization can hold at once.
  • Need more? Contact us and we’ll raise it for your organization.

Treat an OOM like any other cache failure: skip the write and carry on. The addon is billed monthly, with usage metered by data written — see Billing → Addons for current pricing.

Limits

LimitValue
Key lifetime24 hours at most
Write budget1 GiB per rolling 24 hours per organization (can be raised)
Largest key, field or value1 MiB
Largest command8 MiB
Largest reply32 MiB (read bigger collections with HSCAN, SSCAN, LRANGE or GETRANGE)
Simultaneous connections64 per organization
Databases0 only
ProtocolRESP2 and RESP3 (negotiated automatically)
Encryption in transitNone — reachable only inside the platform and over the VPN

Errors

ErrorWhat it meansWhat to do
OOM cache write budget exceeded …You’ve written 1 GiB in the last 24 hours.Skip the write. Reads and deletes still work (deleting doesn’t free budget); writes resume as the window rolls.
ERR value too large …A key, field or value over 1 MiB, or a command over 8 MiB.Store less (split or compress) or skip it. The connection stays usable (only a command of tens of MiB closes it).
ERR max number of clients reachedYour organization already has 64 open connections.Share one client per process; cap the pool (max_connections).
ERR command '…' is not available on the InsightScrap cacheThe command isn’t supported. With evalsha, it’s usually redis-py’s r.lock() releasing.See supported commands; for locks, see Locks.
ERR DB index is out of rangeSELECT with a number other than 0.Use database 0.
ERR GT is not supported …EXPIRE … GT — TTLs can only shrink.Drop GT; set a new lifetime by overwriting the key.
WRONGPASS … / NOAUTH …Missing, wrong or rotated credentials.Connect with the full INSIGHT_CACHE_URL; reveal it again for local use.
ERR cache access was revoked for this credentialThe credential has been revoked.Check the addon is still enabled, then contact us.
ERR reply too large …The answer would be over 32 MiB — usually HGETALL, SMEMBERS or LRANGE 0 -1 on a very large collection. The connection stays usable.Read it in pieces with HSCAN, SSCAN or LRANGE ranges.
ERR cache busy, retry / ERR Protocol error: server busy, retryToo many very large requests or replies at the same moment across the platform. The second form closes the connection.Retry with a short backoff, or treat it as a miss.
ERR cache temporarily unavailableA brief interruption on our side; the connection is closed.Retry, or treat it as a miss.
ERR … inside a transaction / … inside MULTIA connection command (AUTH, HELLO, SELECT, CLIENT, COMMAND, INFO) or FLUSHDB/FLUSHALL inside MULTI (PING and ECHO are fine).Run it outside the transaction.

Supported commands

GroupCommands
Connection & serverAUTH, HELLO, PING, ECHO, QUIT, SELECT 0, CLIENT SETNAME / GETNAME / SETINFO / ID, COMMAND (returns an empty list), INFO (minimal)
Strings & countersGET, SET (EX / PX / EXAT / PXAT / KEEPTTL / NX / XX / GET), SETNX, SETEX, PSETEX, MGET, MSET, MSETNX, GETDEL, GETSET, GETRANGE, STRLEN, APPEND, INCR, INCRBY, INCRBYFLOAT, DECR, DECRBY
KeysDEL, UNLINK, EXISTS, TYPE, TTL, PTTL, EXPIRETIME, PEXPIRETIME, EXPIRE, PEXPIRE, EXPIREAT, PEXPIREAT, TOUCH, SCAN (MATCH / COUNT / TYPE), FLUSHDB / FLUSHALL (delete your organization’s keys only)
HashesHGET, HSET, HSETNX, HMSET, HMGET, HDEL, HEXISTS, HLEN, HSTRLEN, HKEYS, HVALS, HGETALL, HINCRBY, HINCRBYFLOAT, HSCAN
SetsSADD, SREM, SISMEMBER, SMISMEMBER, SCARD, SMEMBERS, SSCAN
ListsLPUSH, RPUSH, LPOP, RPOP, LLEN, LRANGE, LINDEX, LTRIM
Transactions & pipelinesMULTI, EXEC, DISCARD, WATCH, UNWATCH — redis-py’s pipeline() works with its default transaction=True

Not available: Lua scripting and functions (so redis-py’s r.lock() can’t release — see Locks), pub/sub, blocking commands (BLPOP and friends), KEYS (use SCAN), DBSIZE, sorted sets, streams, admin commands (CONFIG, DEBUG, MONITOR, …), PERSIST, RENAME, OBJECT, DUMP / RESTORE. They fail with ERR command '…' is not available on the InsightScrap cache.

Listing keys: use SCAN

KEYS isn’t available — use SCAN. The cache walks a shared keyspace and filters it down to your keys, so a page can come back empty while the cursor is still non-zero. Always loop until the cursor returns 0 (redis-py’s scan_iter does that for you) and ask for big pages with COUNT 1000:

for key in r.scan_iter(match="ean:*", count=1000): print(key)

To wipe everything your organization has stored, use FLUSHDB — it only touches your keys.

FAQ

Can a key disappear before its TTL? Yes. It’s a cache: under memory pressure on the platform, keys may be evicted early. Always treat a miss as normal — fetch the data and write it back.

Is it consistent? A write is visible to every connection in your organization as soon as the command returns. Single commands are atomic (INCR, SET … NX, HSET, …), and MULTI/EXEC with WATCH gives you optimistic transactions.

Does data survive platform maintenance? Treat it as best-effort. The cache isn’t persisted, so a platform restart can empty it. Never keep the only copy of anything here.

Can I keep a key longer than 24 hours? No — PERSIST and longer TTLs aren’t available. Use Storage or the postgres add-on for durable data.

Can other organizations see my keys? No. Keys are private to your organization: you only ever see, SCAN and flush your own.

Can I use it as a job queue, a Celery/RQ broker or for pub/sub? No. Blocking commands, pub/sub, Lua scripts and sorted sets aren’t available, so queue and broker libraries (and distributed schedulers built on them) won’t work. Lists are fine for simple non-blocking push/pop.

Do my local runs share the cache with production jobs? Yes — same organization, same keys. Prefix experimental keys so they don’t mix with real data.

Last updated on