Download historical racing odds in Python: closing lines, price paths and CLV
Live prices tell you what the market thinks now. A model needs what the market thought at the jump, and what happened next. PuntersEdge keeps a closing-line archive for that: for every Australian and New Zealand race a customer-served bookmaker priced, each book's first price captured in the hour before the jump, its last price before the jump, and the result where one is published. This tutorial pulls the archive into pandas, answers two questions with it, draws one runner's price path, and scores bets against the close.
Every block below was run against the live API on 28 September 2026, and the output under each one is what it printed.
18+ only. This is data tooling, not tipping or advice, and nothing below is a staking strategy. Gambling Help: 1800 858 858.
What you need
A paid key. Standard reads the last 90 days of the archive; Plus and above read all of it, and steps 5 and 6 (price paths and closing line value) are Plus and above. Free and Hobby keys get a 403 on the archive and are not charged. See pricing for which plan has what. Set the key as an environment variable and pass it in the X-API-Key header:
import io
import os
import pandas as pd
import requests
API = "https://api.puntersedge.online/v1"
HEADERS = {"X-API-Key": os.environ["PUNTERSEDGE_API_KEY"]}
Step 1: see what the archive holds
/v1/racing/closing-lines/coverage costs 1 credit and describes the whole archive: how far back it goes, how many races and books, and how much of it is a genuine closing line (a last price inside the final five minutes; the rest stopped being quoted early, so their close is a last-seen price).
r = requests.get(f"{API}/racing/closing-lines/coverage", headers=HEADERS, timeout=30)
r.raise_for_status()
cov = r.json()
print(f'{cov["archive_from"][:10]} to {cov["archive_to"][:10]}: {cov["races"]:,} races, '
f'{cov["rows"]:,} rows, {cov["bookmakers"]} bookmakers')
print(f'{cov["is_closing_line_pct"]}% of rows are a genuine closing line')
2026-08-04 to 2026-09-28: 40,289 races, 1,447,227 rows, 14 bookmakers
84.17% of rows are a genuine closing line
Step 2: a day of closing lines into pandas
/v1/racing/closing-lines returns one row per race, runner and bookmaker: open_win_price, close_win_price, is_closing_line, open_is_baseline, finish_position, result_status, the venue and race, and runner_key and runner_ref to join on. A JSON page holds at most 5,000 rows and costs 5 credits, so page through:
def closing_lines(**params):
"""Every row for a query. JSON pages hold at most 5,000 rows, so page through."""
rows, offset = [], 0
while True:
r = requests.get(f"{API}/racing/closing-lines", headers=HEADERS, timeout=120,
params={**params, "limit": 5000, "offset": offset})
r.raise_for_status()
page = r.json()
rows += page["rows"]
offset += page["rows_returned"]
if page["rows_returned"] == 0 or offset >= page["total_rows"]:
return pd.DataFrame(rows)
df = closing_lines(date="2026-09-26", country="AU", categories="horse",
resulted_only="true")
print(f"{len(df):,} rows: {df.race_id.nunique()} races, "
f"{df.runner_key.nunique()} runners, {df.bookmaker_key.nunique()} bookmakers")
7,630 rows: 80 races, 784 runners, 12 bookmakers
For bulk pulls, format=csv returns up to 50,000 rows in one request for 20 credits, with the same columns:
r = requests.get(f"{API}/racing/closing-lines", headers=HEADERS, timeout=120, params={
"date": "2026-09-26", "country": "AU", "categories": "horse",
"resulted_only": "true", "format": "csv", "limit": 50000})
r.raise_for_status()
df_csv = pd.read_csv(io.StringIO(r.text))
print(f"{len(df_csv):,} rows, {len(df_csv.columns)} columns")
7,630 rows, 29 columns
Step 3: how often did the favourite win?
Take each runner's best closing price across the books, call the shortest runner in each race the favourite, and check it against the result. Two details matter. Group on runner_key, not runner_name: books spell names differently, so one horse can appear as UPTOWN MONK and Uptown Monk. And results publish placegetters only, so an unplaced runner has no finish_position; in a final result that means it did not win, which is what eq(1) gives you.
closes = df[df.is_closing_line & df.result_status.eq("final")]
best = (closes.groupby(["race_id", "runner_key"], as_index=False)
.agg(venue=("venue", "first"), race_number=("race_number", "first"),
runner_name=("runner_name", "first"),
finish_position=("finish_position", "max"),
close=("close_win_price", "max")))
fav = best.loc[best.groupby("race_id")["close"].idxmin()]
won = fav.finish_position.eq(1)
print(f"The favourite won {won.sum()} of {len(fav)} races ({won.mean():.0%}), "
f"median price ${fav.close.median():.2f}")
print(f"$1 on every favourite at the best closing price: "
f"{(fav.close * won).sum() - len(fav):+.2f} units")
The favourite won 32 of 80 races (40%), median price $2.90
$1 on every favourite at the best closing price: +2.95 units
One day is 80 races, which is noise, and the best closing price across a dozen books is not a price anyone reliably gets at the jump. The point is the method: the same code over 90 days (Standard) or the whole archive (Plus) gives you a real baseline to measure a model against.
Step 4: did late market moves beat their price?
Here open_win_price is the first price captured in the final hour, and open_is_baseline marks the series that start at the beginning of that hour, so keep only those. Bucket each runner by how its price moved in the last hour, then compare the wins each bucket actually had with the wins its closing prices implied (the sum of 1/price):
week = closing_lines(**{"from": "2026-09-20", "to": "2026-09-27"},
country="AU", categories="horse", resulted_only="true")
w = week[week.is_closing_line & week.open_is_baseline & week.result_status.eq("final")].copy()
w["move"] = w.close_win_price / w.open_win_price - 1
runners = (w.groupby(["race_id", "runner_key"], as_index=False)
.agg(move=("move", "median"),
won=("finish_position", lambda s: s.eq(1).any()),
close=("close_win_price", "median")))
runners["bucket"] = pd.cut(runners.move, [-1, -0.2, -0.05, 0.05, 0.2, 99],
labels=["shortened 20%+", "shortened 5-20%", "steady",
"drifted 5-20%", "drifted 20%+"])
table = runners.groupby("bucket", observed=True).agg(
runners=("won", "size"), wins=("won", "sum"),
implied_wins=("close", lambda p: (1 / p).sum()))
table["wins / implied"] = table.wins / table.implied_wins
print(f"{runners.race_id.nunique()} races, {len(runners):,} runners")
print(table.round(2).to_string())
345 races, 3,336 runners
runners wins implied_wins wins / implied
bucket
shortened 20%+ 296 23 33.88 0.68
shortened 5-20% 658 74 111.12 0.67
steady 568 91 96.53 0.94
drifted 5-20% 736 95 102.38 0.93
drifted 20%+ 1078 61 71.83 0.85
Every bucket sits below 1.0 because a bookmaker's prices carry a margin. What a model looks for is the relative gap: in this week, runners that shortened in the last hour won about two-thirds of what their closing price implied, while steady and drifting runners came closer to their price. It is one week and 345 races, so run it over the full window before believing any of it.
Step 5: one runner's price path
/v1/racing/price-paths returns every captured price point, per runner and book, from the start of the window to the last pre-jump quote (5 credits, Plus and above; on Standard, /v1/racing/price-history answers the same question for one race). Here is the first favourite from step 3, book by book:
race = fav.iloc[0]
r = requests.get(f"{API}/racing/price-paths", headers=HEADERS, timeout=60,
params={"race_id": race.race_id, "limit": 5000})
r.raise_for_status()
path = pd.DataFrame(r.json()["rows"])
mine = path[path.runner_key == race.runner_key].sort_values("captured_at")
print(f"{race.venue} R{race.race_number}, {race.runner_name}: "
f"{len(mine)} price points from {mine.bookmaker_key.nunique()} books")
print(mine.groupby("bookmaker_key")["win_price"].agg(["first", "last", "count"]).head(4).to_string())
Murray Bridge R4, Uptown Monk: 126 price points from 13 books
first last count
bookmaker_key
betgold 3.6 3.1 12
betr_au 3.7 3.1 7
betright 3.7 3.2 10
boostbet 3.6 3.1 12
Step 6: score your own bets against the close
POST /v1/racing/clv takes up to 200 bets per request for 5 credits (Plus and above). Name each race by race_id, or by venue, meeting date and race number, and the runner by runner_ref, saddlecloth or name. Each bet comes back with the closing price it is measured against (the best close by default, "reference": "median" for the median, or the named bookmaker's own close if you send one), its closing line value, whether it beat the close, and once the result is final, whether it won and its profit or loss:
bets = [
{"ref": "bet-1", "race_id": fav.iloc[0].race_id, "runner_name": fav.iloc[0].runner_name,
"price": 3.75, "stake": 10},
{"ref": "bet-2", "race_id": fav.iloc[1].race_id, "runner_name": fav.iloc[1].runner_name,
"price": 3.10, "stake": 10},
]
r = requests.post(f"{API}/racing/clv", headers=HEADERS, timeout=60, json={"bets": bets})
r.raise_for_status()
res = r.json()
for b in res["bets"]:
print(f'{b["ref"]}: took {b["price_taken"]}, closed {b["close_reference"]}, '
f'CLV {b["clv_pct"]:+.1f}%, won {b["won"]}, P&L {b["pnl"]:+.2f}')
s = res["summary"]
print(f'average CLV {s["avg_clv_pct"]:+.2f}%, beat the close {s["beat_close"]} of {s["scored"]}')
bet-1: took 3.75, closed 3.4, CLV +10.3%, won False, P&L -10.00
bet-2: took 3.1, closed 3.3, CLV -6.1%, won False, P&L -10.00
average CLV +2.11%, beat the close 1 of 2
Beating the close is the signal a model is finding value before the market does; winning or losing one bet is not. Across hundreds of bets, the average closing line value is the number to watch.
Things that will bite you
- Runners the result marked scratched are hidden from
/v1/racing/closing-linesby default; passinclude_scratched=trueto see them. Price paths keep them, with the flag. - Rows flagged as a possible split meeting or a runner-name fragment are left out unless you pass
include_flagged=true. - Results are collected for Australian and New Zealand racing only, from 15 August 2026. Races elsewhere carry closing prices but never a finishing position.
- BetDeluxe closes in the archive start on 28 September 2026. The data quality page shows the archive's own quality flags, recomputed from production.