Clusters

A cluster is a saved list of websites. Through the API you can make one from a search or from your own list, combine clusters, and pull values out of the page source of every site in one - contacts, social profiles, analytics and tag ids, anything a regular expression can match.

The endpoints

EndpointMethodWhat it does
/v1/clustersGETThe account's clusters, newest first: id, name, number of domains, creation time.
/v1/clustersPOSTCreate a cluster from a search (query) or from a list (domains); optional name.
/v1/clusters/{id}GETThe cluster and a page of its domains: page, per_page (up to 10000), format=txt for one domain per line.
/v1/clusters/combinePOSTA new cluster from existing ones: operation and, or or diff, clusters - a list of ids.
/v1/clusters/{id}/extractGET, POSTExtract values with presets and regex, part by part: offset, limit (up to 1000 sites per call), format json, xml or csv.
/v1/clusters/presetsGETThe ready-made expressions with their patterns.
/v1/clusters/{id}/renamePOSTNew name.
/v1/clusters/{id}/deletePOSTDelete the cluster for good.

Everything that changes a cluster is a POST with a JSON body - no PUT or DELETE, so any client that can search can manage clusters too. The token, the rate limit and the error format are the same as for search; the cluster numbers in your /v1/account answer show how many you have and how many extraction points are left today.

Creating a cluster

From a search: the results of the query, up to your plan's rows per report, become the cluster. It costs one search, the same as /v1/search.

curl https://api.publicwww.com/v1/clusters \
     -H "Authorization: Bearer $PUBLICWWW_KEY" \
     -H "Content-Type: application/json" \
     -d '{"query": "\"googletagmanager.com/gtm.js\"", "name": "GTM sites"}'

{"id": 7, "name": "GTM sites", "size": 100000, "created": "2026-10-10T20:21:30Z",
 "query": "\"googletagmanager.com/gtm.js\"", "total": 2412577, "index_complete": true}

From a list: domains or urls, as a JSON array or one per line. Only sites that are in the index are kept - submitted says how many you sent, size how many made it. A list costs nothing.

curl https://api.publicwww.com/v1/clusters \
     -H "Authorization: Bearer $PUBLICWWW_KEY" \
     -H "Content-Type: application/json" \
     -d '{"domains": ["example.com", "https://www.example.org/about"], "name": "Prospects"}'

The plan caps the number of domains in a cluster, and an account keeps up to 100 clusters. With 100 already there, creating another answers 409 cluster_limit - nothing is deleted on your behalf; delete the ones you no longer need first.

Combining clusters

curl https://api.publicwww.com/v1/clusters/combine \
     -H "Authorization: Bearer $PUBLICWWW_KEY" \
     -H "Content-Type: application/json" \
     -d '{"operation": "diff", "clusters": [7, 3], "name": "GTM, not yet contacted"}'

and keeps the domains that are in every cluster, or those in any of them, and diff those in the first cluster but not in the second - exactly two. Combining costs nothing.

Extracting data

Extraction reads the indexed pages of every site in the cluster and returns what the expressions capture, one column per expression. The easiest way is a preset:

PresetWhat it pulls out
emailAddresses from mailto: links
phoneNumbers from tel: links
whatsapp, telegram, skypeWhatsApp numbers, Telegram usernames, Skype names from their links
facebook, instagram, twitter, linkedinLinks to the site's social profiles (twitter also takes x.com, linkedin company pages)
gtm, ga4, uaGoogle Tag Manager container, Google Analytics 4 and Universal Analytics ids
hotjarHotjar site id
adsenseAdSense publisher id
bitcoinAddresses from bitcoin: payment links

Or write your own: a regular expression in slashes (or pipes) with optional i, m, s, u flags, up to 200 characters. The first capturing group is the value - the same rule as snipexp:. Up to ten expressions per call, presets included.

curl https://api.publicwww.com/v1/clusters/7/extract \
     -H "Authorization: Bearer $PUBLICWWW_KEY" \
     -H "Content-Type: application/json" \
     -d '{"presets": ["gtm", "email"], "regex": ["/data-site-id=\"([0-9]+)\"/i"], "limit": 1000}'

{
  "cluster": 7, "name": "GTM sites", "size": 100000,
  "offset": 0, "scanned": 1000, "in_index": 1000, "with_matches": 941,
  "next_offset": 1000,
  "regex": ["/(GTM-[A-Z0-9]{4,10})\\b/", "/mailto:(...)/i", "/data-site-id=\"([0-9]+)\"/i"],
  "points_used": 1834.2, "points_left": 98165,
  "rows": [
    { "domain": "example.com", "values": [["GTM-AB12CD"], ["info@example.com"], []], "matches": 2 }
  ]
}

Part by part. One call goes through up to 1000 sites of the cluster, starting at offset. Call again with offset set to next_offset until it comes back null. Sites where nothing matched are left out; skip_empty=0 lists them too. format=csv gives a row per site - the domain, then a column per expression - with the next offset in the X-Next-Offset header.

Points. Extraction spends the daily extraction allowance of your plan: each site found in the index costs as many points as values were found on it, 0.1 if none. When the points run out the call stops early with "stopped": "extract_quota_exceeded" and returns what it has; a call with nothing left gets 429 extract_quota_exceeded. Try an expression on a small limit first - see cluster limits.

A whole cluster, in code

Create a cluster from a search and write the extracted values of every site into a CSV file. The client libraries have the same as ready functions and a command line.

Python

import csv, os, time, requests

KEY  = os.environ["PUBLICWWW_KEY"]
BASE = "https://api.publicwww.com"
H    = {"Authorization": "Bearer " + KEY}

def call(method, path, body=None):
    while True:
        r = requests.request(method, BASE + path, headers=H, json=body)
        if r.status_code == 429 and r.json()["error"]["code"] == "too_many_requests":
            time.sleep(int(r.headers.get("Retry-After", 30)))
            continue
        r.raise_for_status()
        return r.json()

cluster = call("POST", "/v1/clusters", {"query": '"googletagmanager.com/gtm.js"'})
offset = 0
with open("extract.csv", "w", newline="") as f:
    out = csv.writer(f)
    while offset is not None:
        part = call("POST", "/v1/clusters/%d/extract" % cluster["id"],
                    {"presets": ["gtm", "email"], "offset": offset})
        for row in part["rows"]:
            out.writerow([row["domain"]] + [" ".join(v) for v in row["values"]])
        offset = part["next_offset"]

JavaScript (Node 18+)

const BASE = "https://api.publicwww.com";
const H = { Authorization: "Bearer " + process.env.PUBLICWWW_KEY,
            "Content-Type": "application/json" };

async function call(method, path, body) {
  for (;;) {
    const r = await fetch(BASE + path, { method, headers: H,
                                         body: body && JSON.stringify(body) });
    const data = await r.json();
    if (r.status === 429 && data.error.code === "too_many_requests") {
      await new Promise(ok => setTimeout(ok, 1000 * (r.headers.get("Retry-After") || 30)));
      continue;
    }
    if (!r.ok) throw new Error(data.error.message);
    return data;
  }
}

const cluster = await call("POST", "/v1/clusters", { query: '"hotjar.com"' });
for (let offset = 0; offset !== null; ) {
  const part = await call("POST", `/v1/clusters/${cluster.id}/extract`,
                          { presets: ["hotjar", "email"], offset });
  for (const row of part.rows) console.log(row.domain, row.values.map(v => v.join(" ")).join(";"));
  offset = part.next_offset;
}

PHP

<?php
function call ($method, $path, $body = null) {
    $ch = curl_init ("https://api.publicwww.com" . $path);
    curl_setopt_array ($ch, [
        CURLOPT_CUSTOMREQUEST  => $method,
        CURLOPT_RETURNTRANSFER => true,
        CURLOPT_HTTPHEADER     => ["Authorization: Bearer " . getenv ("PUBLICWWW_KEY"),
                                   "Content-Type: application/json"],
        CURLOPT_POSTFIELDS     => $body === null ? null : json_encode ($body),
    ]);
    $data = json_decode (curl_exec ($ch), true);
    if (isset ($data ["error"])) throw new Exception ($data ["error"]["message"]);
    return $data;
}

$cluster = call ("POST", "/v1/clusters", ["query" => '"jquery.min.js"']);
$out = fopen ("extract.csv", "w");
for ($offset = 0; $offset !== null; ) {
    $part = call ("POST", "/v1/clusters/" . $cluster ["id"] . "/extract",
                  ["presets" => ["email", "phone"], "offset" => $offset]);
    foreach ($part ["rows"] as $row)
        fputcsv ($out, array_merge ([$row ["domain"]], array_map (fn ($v) => join (" ", $v), $row ["values"])));
    $offset = $part ["next_offset"];
}

Go

package main

import (
	"bytes"
	"encoding/json"
	"fmt"
	"net/http"
	"os"
	"strings"
)

func call(method, path string, body, out any) error {
	b, _ := json.Marshal(body)
	req, _ := http.NewRequest(method, "https://api.publicwww.com"+path, bytes.NewReader(b))
	req.Header.Set("Authorization", "Bearer "+os.Getenv("PUBLICWWW_KEY"))
	req.Header.Set("Content-Type", "application/json")
	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		return err
	}
	defer resp.Body.Close()
	if resp.StatusCode >= 300 {
		return fmt.Errorf("publicwww: %s", resp.Status)
	}
	return json.NewDecoder(resp.Body).Decode(out)
}

func main() {
	var cluster struct{ ID int `json:"id"` }
	if err := call("POST", "/v1/clusters", map[string]any{"query": `"googletagmanager.com/gtm.js"`}, &cluster); err != nil {
		panic(err)
	}
	for offset := 0; ; {
		var part struct {
			Rows []struct {
				Domain string     `json:"domain"`
				Values [][]string `json:"values"`
			} `json:"rows"`
			NextOffset *int `json:"next_offset"`
		}
		path := fmt.Sprintf("/v1/clusters/%d/extract", cluster.ID)
		if err := call("POST", path, map[string]any{"presets": []string{"gtm", "ga4"}, "offset": offset}, &part); err != nil {
			panic(err)
		}
		for _, r := range part.Rows {
			cells := []string{r.Domain}
			for _, v := range r.Values {
				cells = append(cells, strings.Join(v, " "))
			}
			fmt.Println(strings.Join(cells, ";"))
		}
		if part.NextOffset == nil {
			break
		}
		offset = *part.NextOffset
	}
}

Ruby

require "json"
require "net/http"

def call(path, body)
  uri = URI("https://api.publicwww.com" + path)
  req = Net::HTTP::Post.new(uri, "Authorization" => "Bearer #{ENV.fetch('PUBLICWWW_KEY')}",
                                  "Content-Type" => "application/json")
  req.body = body.to_json
  res = Net::HTTP.start(uri.host, uri.port, use_ssl: true) { |h| h.request(req) }
  data = JSON.parse(res.body)
  raise data["error"]["message"] if data["error"]
  data
end

cluster = call("/v1/clusters", { query: '"hotjar.com"' })
offset = 0
while offset
  part = call("/v1/clusters/#{cluster['id']}/extract", { presets: %w[hotjar email], offset: offset })
  part["rows"].each { |r| puts [r["domain"], *r["values"].map { |v| v.join(" ") }].join(";") }
  offset = part["next_offset"]
end

Errors

Status and codeMeaning
404 cluster_not_foundNo cluster with that id on this account.
409 cluster_limitThe account already keeps 100 clusters.
400 invalid_regexAn expression is not a PCRE in slashes or pipes, up to 200 characters.
400 unknown_presetNo such preset; the answer lists them.
400 missing_source, ambiguous_sourceCreating needs either query or domains.
429 extract_quota_exceededToday's extraction points are used up.

The full list is on the errors page and in the API's own description at https://api.publicwww.com/. In an AI assistant the same operations are MCP tools.

Next Code examples