Clusters
A cluster is a saved list of websites. Through the API you can make one from a search or from your own list, combine clusters, and pull values out of the page source of every site in one - contacts, social profiles, analytics and tag ids, anything a regular expression can match.
The endpoints
| Endpoint | Method | What it does |
|---|---|---|
/v1/clusters | GET | The account's clusters, newest first: id, name, number of domains, creation time. |
/v1/clusters | POST | Create a cluster from a search (query) or from a list (domains); optional name. |
/v1/clusters/{id} | GET | The cluster and a page of its domains: page, per_page (up to 10000), format=txt for one domain per line. |
/v1/clusters/combine | POST | A new cluster from existing ones: operation and, or or diff, clusters - a list of ids. |
/v1/clusters/{id}/extract | GET, POST | Extract values with presets and regex, part by part: offset, limit (up to 1000 sites per call), format json, xml or csv. |
/v1/clusters/presets | GET | The ready-made expressions with their patterns. |
/v1/clusters/{id}/rename | POST | New name. |
/v1/clusters/{id}/delete | POST | Delete the cluster for good. |
Everything that changes a cluster is a POST with a JSON body - no PUT or
DELETE, so any client that can search can manage clusters too. The token,
the rate limit and the error format are the same as for
search; the cluster numbers in your
/v1/account answer show how many you have and how many
extraction points are left today.
Creating a cluster
From a search: the results of the query, up to your plan's rows per
report, become the cluster. It costs one search, the same as
/v1/search.
curl https://api.publicwww.com/v1/clusters \
-H "Authorization: Bearer $PUBLICWWW_KEY" \
-H "Content-Type: application/json" \
-d '{"query": "\"googletagmanager.com/gtm.js\"", "name": "GTM sites"}'
{"id": 7, "name": "GTM sites", "size": 100000, "created": "2026-10-10T20:21:30Z",
"query": "\"googletagmanager.com/gtm.js\"", "total": 2412577, "index_complete": true}
From a list: domains or urls, as a JSON array or one per line. Only
sites that are in the index are kept - submitted says how many
you sent, size how many made it. A list costs nothing.
curl https://api.publicwww.com/v1/clusters \
-H "Authorization: Bearer $PUBLICWWW_KEY" \
-H "Content-Type: application/json" \
-d '{"domains": ["example.com", "https://www.example.org/about"], "name": "Prospects"}'
The plan caps the number of domains in a cluster, and an account keeps up to
100 clusters. With 100 already there, creating another answers
409 cluster_limit - nothing is deleted on your behalf; delete
the ones you no longer need first.
Combining clusters
curl https://api.publicwww.com/v1/clusters/combine \
-H "Authorization: Bearer $PUBLICWWW_KEY" \
-H "Content-Type: application/json" \
-d '{"operation": "diff", "clusters": [7, 3], "name": "GTM, not yet contacted"}'
and keeps the domains that are in every cluster, or
those in any of them, and diff those in the first cluster but not
in the second - exactly two. Combining costs nothing.
Extracting data
Extraction reads the indexed pages of every site in the cluster and returns what the expressions capture, one column per expression. The easiest way is a preset:
| Preset | What it pulls out |
|---|---|
email | Addresses from mailto: links |
phone | Numbers from tel: links |
whatsapp, telegram, skype | WhatsApp numbers, Telegram usernames, Skype names from their links |
facebook, instagram, twitter, linkedin | Links to the site's social profiles (twitter also takes x.com, linkedin company pages) |
gtm, ga4, ua | Google Tag Manager container, Google Analytics 4 and Universal Analytics ids |
hotjar | Hotjar site id |
adsense | AdSense publisher id |
bitcoin | Addresses from bitcoin: payment links |
Or write your own: a regular expression in slashes (or pipes) with optional
i, m, s, u flags, up to 200
characters. The first capturing group is the value - the same rule as
snipexp:. Up to ten
expressions per call, presets included.
curl https://api.publicwww.com/v1/clusters/7/extract \
-H "Authorization: Bearer $PUBLICWWW_KEY" \
-H "Content-Type: application/json" \
-d '{"presets": ["gtm", "email"], "regex": ["/data-site-id=\"([0-9]+)\"/i"], "limit": 1000}'
{
"cluster": 7, "name": "GTM sites", "size": 100000,
"offset": 0, "scanned": 1000, "in_index": 1000, "with_matches": 941,
"next_offset": 1000,
"regex": ["/(GTM-[A-Z0-9]{4,10})\\b/", "/mailto:(...)/i", "/data-site-id=\"([0-9]+)\"/i"],
"points_used": 1834.2, "points_left": 98165,
"rows": [
{ "domain": "example.com", "values": [["GTM-AB12CD"], ["info@example.com"], []], "matches": 2 }
]
}
Part by part. One call goes through up to 1000 sites of the cluster,
starting at offset. Call again with offset set to
next_offset until it comes back null. Sites where
nothing matched are left out; skip_empty=0 lists them too.
format=csv gives a row per site - the domain, then a column per
expression - with the next offset in the X-Next-Offset header.
Points. Extraction spends the daily extraction allowance of your
plan: each site found in the index costs as many points as values were found
on it, 0.1 if none. When the points run out the call stops early with
"stopped": "extract_quota_exceeded" and returns what it has; a
call with nothing left gets 429 extract_quota_exceeded. Try an
expression on a small limit first - see
cluster limits.
A whole cluster, in code
Create a cluster from a search and write the extracted values of every site into a CSV file. The client libraries have the same as ready functions and a command line.
Python
import csv, os, time, requests
KEY = os.environ["PUBLICWWW_KEY"]
BASE = "https://api.publicwww.com"
H = {"Authorization": "Bearer " + KEY}
def call(method, path, body=None):
while True:
r = requests.request(method, BASE + path, headers=H, json=body)
if r.status_code == 429 and r.json()["error"]["code"] == "too_many_requests":
time.sleep(int(r.headers.get("Retry-After", 30)))
continue
r.raise_for_status()
return r.json()
cluster = call("POST", "/v1/clusters", {"query": '"googletagmanager.com/gtm.js"'})
offset = 0
with open("extract.csv", "w", newline="") as f:
out = csv.writer(f)
while offset is not None:
part = call("POST", "/v1/clusters/%d/extract" % cluster["id"],
{"presets": ["gtm", "email"], "offset": offset})
for row in part["rows"]:
out.writerow([row["domain"]] + [" ".join(v) for v in row["values"]])
offset = part["next_offset"]
JavaScript (Node 18+)
const BASE = "https://api.publicwww.com";
const H = { Authorization: "Bearer " + process.env.PUBLICWWW_KEY,
"Content-Type": "application/json" };
async function call(method, path, body) {
for (;;) {
const r = await fetch(BASE + path, { method, headers: H,
body: body && JSON.stringify(body) });
const data = await r.json();
if (r.status === 429 && data.error.code === "too_many_requests") {
await new Promise(ok => setTimeout(ok, 1000 * (r.headers.get("Retry-After") || 30)));
continue;
}
if (!r.ok) throw new Error(data.error.message);
return data;
}
}
const cluster = await call("POST", "/v1/clusters", { query: '"hotjar.com"' });
for (let offset = 0; offset !== null; ) {
const part = await call("POST", `/v1/clusters/${cluster.id}/extract`,
{ presets: ["hotjar", "email"], offset });
for (const row of part.rows) console.log(row.domain, row.values.map(v => v.join(" ")).join(";"));
offset = part.next_offset;
}
PHP
<?php
function call ($method, $path, $body = null) {
$ch = curl_init ("https://api.publicwww.com" . $path);
curl_setopt_array ($ch, [
CURLOPT_CUSTOMREQUEST => $method,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => ["Authorization: Bearer " . getenv ("PUBLICWWW_KEY"),
"Content-Type: application/json"],
CURLOPT_POSTFIELDS => $body === null ? null : json_encode ($body),
]);
$data = json_decode (curl_exec ($ch), true);
if (isset ($data ["error"])) throw new Exception ($data ["error"]["message"]);
return $data;
}
$cluster = call ("POST", "/v1/clusters", ["query" => '"jquery.min.js"']);
$out = fopen ("extract.csv", "w");
for ($offset = 0; $offset !== null; ) {
$part = call ("POST", "/v1/clusters/" . $cluster ["id"] . "/extract",
["presets" => ["email", "phone"], "offset" => $offset]);
foreach ($part ["rows"] as $row)
fputcsv ($out, array_merge ([$row ["domain"]], array_map (fn ($v) => join (" ", $v), $row ["values"])));
$offset = $part ["next_offset"];
}
Go
package main
import (
"bytes"
"encoding/json"
"fmt"
"net/http"
"os"
"strings"
)
func call(method, path string, body, out any) error {
b, _ := json.Marshal(body)
req, _ := http.NewRequest(method, "https://api.publicwww.com"+path, bytes.NewReader(b))
req.Header.Set("Authorization", "Bearer "+os.Getenv("PUBLICWWW_KEY"))
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
if resp.StatusCode >= 300 {
return fmt.Errorf("publicwww: %s", resp.Status)
}
return json.NewDecoder(resp.Body).Decode(out)
}
func main() {
var cluster struct{ ID int `json:"id"` }
if err := call("POST", "/v1/clusters", map[string]any{"query": `"googletagmanager.com/gtm.js"`}, &cluster); err != nil {
panic(err)
}
for offset := 0; ; {
var part struct {
Rows []struct {
Domain string `json:"domain"`
Values [][]string `json:"values"`
} `json:"rows"`
NextOffset *int `json:"next_offset"`
}
path := fmt.Sprintf("/v1/clusters/%d/extract", cluster.ID)
if err := call("POST", path, map[string]any{"presets": []string{"gtm", "ga4"}, "offset": offset}, &part); err != nil {
panic(err)
}
for _, r := range part.Rows {
cells := []string{r.Domain}
for _, v := range r.Values {
cells = append(cells, strings.Join(v, " "))
}
fmt.Println(strings.Join(cells, ";"))
}
if part.NextOffset == nil {
break
}
offset = *part.NextOffset
}
}
Ruby
require "json"
require "net/http"
def call(path, body)
uri = URI("https://api.publicwww.com" + path)
req = Net::HTTP::Post.new(uri, "Authorization" => "Bearer #{ENV.fetch('PUBLICWWW_KEY')}",
"Content-Type" => "application/json")
req.body = body.to_json
res = Net::HTTP.start(uri.host, uri.port, use_ssl: true) { |h| h.request(req) }
data = JSON.parse(res.body)
raise data["error"]["message"] if data["error"]
data
end
cluster = call("/v1/clusters", { query: '"hotjar.com"' })
offset = 0
while offset
part = call("/v1/clusters/#{cluster['id']}/extract", { presets: %w[hotjar email], offset: offset })
part["rows"].each { |r| puts [r["domain"], *r["values"].map { |v| v.join(" ") }].join(";") }
offset = part["next_offset"]
end
Errors
| Status and code | Meaning |
|---|---|
404 cluster_not_found | No cluster with that id on this account. |
409 cluster_limit | The account already keeps 100 clusters. |
400 invalid_regex | An expression is not a PCRE in slashes or pipes, up to 200 characters. |
400 unknown_preset | No such preset; the answer lists them. |
400 missing_source, ambiguous_source | Creating needs either query or domains. |
429 extract_quota_exceeded | Today's extraction points are used up. |
The full list is on the errors page and in the API's own description at https://api.publicwww.com/. In an AI assistant the same operations are MCP tools.
Next Code examples