0

Research company domains with Apify and review unresolved matches

by
Published yesterday

Turn 1–10 public company domains into company research records using LinkedIn Company by Domain on Apify. Maintained by the developer of the linked actor. ## Setup 1. Bring your own Apify account and check [current actor pricing](https://apify.com/george.the.developer/linkedin-company-by-domain). 2. Save your private token in an `apify_api_key` resource with an `api_key` field; select it for `credentials`. Never paste a token into source or publish resource values. 3. Set the script timeout to 20 minutes (1200 seconds), where your instance permits it. Keep automatic retries and perpetual execution disabled; disable retries for this step if used in a flow. 4. Pass a list such as `["stripe.com", "vercel.com", "gitlab.com"]` to `domains`. Run deliberately: Apify charges and Windmill usage are separate. ## Behavior and limits The script starts one asynchronous actor run with 512 MB memory, a 900-second timeout and a $0.10 maximum actor charge per invocation. It polls the same run and retrieves up to 100 rows only after SUCCEEDED. The cap and timeout can leave domains missing or unresolved. A failed or expired request does not abort the actor. Inspect Apify Console before retrying an ambiguous start; each new invocation can create another paid run. Review `rows`, `missingDomains` and `needsReview`. Unresolved rows are retained. Parent-domain matches are flagged rather than assumed to prove the requested subdomain identity. Field screening does not independently verify a company. This returns company records, not employee profiles or email addresses. ## Verification Thirteen Python tests exercise this script against a local HTTP server with synthetic credentials. These checks do not establish native Windmill execution or live actor resolution. No paid run was used for this submission. [Source and setup guide](https://the-ai-entrepreneur-ai-hub.github.io/apify-company-research-workflows/guides/windmill.html). This pack is not endorsed by Windmill.

Script apify
  • Submitted by john.haley.81.front.head531 Python3
    Created 2 days ago
    1
    """Company research for Windmill; uses the caller's Apify resource."""
    2
    
    
    3
    import ipaddress
    4
    import json
    5
    import re
    6
    import time
    7
    from typing import TypedDict
    8
    from urllib.error import HTTPError, URLError
    9
    from urllib.parse import urlsplit
    10
    from urllib.request import HTTPRedirectHandler, Request, build_opener
    11
    
    
    12
    API_BASE = 'https://api.apify.com/v2'
    13
    ACTIVE = {'READY', 'RUNNING', 'ABORTING', 'TIMING-OUT'}
    14
    TERMINAL = {'SUCCEEDED', 'FAILED', 'ABORTED', 'TIMED-OUT'}
    15
    
    
    16
    
    
    17
    class apify_api_key(TypedDict):
    18
        api_key: str
    19
    
    
    20
    
    
    21
    class _NoRedirect(HTTPRedirectHandler):
    22
        def redirect_request(self, req, fp, code, msg, headers, newurl):
    23
            return None
    24
    
    
    25
    
    
    26
    def _domain(value):
    27
        if not isinstance(value, str) or not value.strip():
    28
            raise ValueError('Each domain must be a nonempty domain or HTTP(S) URL.')
    29
        value = value.strip()
    30
        if any(char.isspace() or ord(char) < 32 for char in value):
    31
            raise ValueError('Domains and URLs cannot contain whitespace or control characters.')
    32
        try:
    33
            parsed = urlsplit(value if '://' in value else 'https://' + value)
    34
            if (parsed.scheme not in ('http', 'https') or parsed.username is not None
    35
                    or parsed.password is not None or parsed.port is not None):
    36
                raise ValueError()
    37
            host = (parsed.hostname or '').encode('idna').decode('ascii').lower()
    38
            if host.startswith('www.'):
    39
                host = host[4:]
    40
            try:
    41
                ipaddress.ip_address(host)
    42
            except ValueError:
    43
                pass
    44
            else:
    45
                raise ValueError()
    46
            labels = host.split('.')
    47
            if (len(host) > 253 or len(labels) < 2 or labels[-1].isdigit()
    48
                    or any(not re.fullmatch(r'[a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?', label)
    49
                           for label in labels)):
    50
                raise ValueError()
    51
            return host
    52
        except (ValueError, UnicodeError):
    53
            raise ValueError('Each domain must be a public domain without credentials or a port.') from None
    54
    
    
    55
    
    
    56
    def _id(value):
    57
        return isinstance(value, str) and re.fullmatch(r'[A-Za-z0-9]{1,64}', value) is not None
    58
    
    
    59
    
    
    60
    def _linkedin(value):
    61
        if not isinstance(value, str):
    62
            return False
    63
        try:
    64
            url = urlsplit(value)
    65
            return (url.scheme == 'https' and url.hostname in ('linkedin.com', 'www.linkedin.com')
    66
                    and url.username is None and url.password is None and url.port is None
    67
                    and re.fullmatch(r'/company/[^/\s]+/?', url.path) is not None)
    68
        except ValueError:
    69
            return False
    70
    
    
    71
    
    
    72
    def _request(opener, token, path, body=None):
    73
        request = Request(API_BASE + path, data=json.dumps(body).encode() if body is not None else None,
    74
                          headers={'Authorization': 'Bearer ' + token,
    75
                                   'Content-Type': 'application/json',
    76
                                   'x-apify-integration-platform': 'windmill'})
    77
        try:
    78
            with opener.open(request, timeout=30) as response:
    79
                return json.load(response)
    80
        except HTTPError as error:
    81
            error.close()
    82
            raise RuntimeError('Apify request failed.') from None
    83
        except (URLError, OSError, ValueError):
    84
            raise RuntimeError('Apify request failed or returned invalid JSON.') from None
    85
    
    
    86
    
    
    87
    def _run(response):
    88
        data = response.get('data') if isinstance(response, dict) else None
    89
        if (not isinstance(data, dict) or not _id(data.get('id'))
    90
                or not isinstance(data.get('status'), str)
    91
                or data.get('status') not in ACTIVE | TERMINAL):
    92
            raise RuntimeError('Invalid actor run response.')
    93
        return data
    94
    
    
    95
    
    
    96
    def main(credentials: apify_api_key, domains: list[str]) -> dict:
    97
        """Start once, poll the same run, and return company rows plus review flags."""
    98
        token = credentials.get('api_key') if isinstance(credentials, dict) else None
    99
        if not isinstance(token, str) or not token or any(ord(c) < 33 or ord(c) > 126 for c in token):
    100
            raise ValueError('Select your own apify_api_key resource with a valid api_key.')
    101
        if not isinstance(domains, list) or not 1 <= len(domains) <= 10:
    102
            raise ValueError('Supply between 1 and 10 domains.')
    103
        requested = list(dict.fromkeys(_domain(value) for value in domains))
    104
        opener = build_opener(_NoRedirect())
    105
        try:
    106
            run = _run(_request(opener, token,
    107
                '/acts/george.the.developer~linkedin-company-by-domain/runs'
    108
                '?memory=512&timeout=900&maxTotalChargeUsd=0.1',
    109
                {'domains': requested, 'maxDomains': len(requested),
    110
                 'mode': 'resolve', 'includeUnresolved': True}))
    111
        except RuntimeError:
    112
            raise RuntimeError('Start was not confirmed. Inspect Apify Console before retrying; '
    113
                               'a paid run may already exist.') from None
    114
        run_id = run['id']
    115
        try:
    116
            deadline = time.monotonic() + 17 * 60
    117
            while run['status'] in ACTIVE:
    118
                if time.monotonic() >= deadline:
    119
                    raise RuntimeError('Polling expired; inspect the existing run in Console.')
    120
                run = _run(_request(opener, token, '/actor-runs/' + run_id))
    121
                if run['id'] != run_id:
    122
                    raise RuntimeError('Run response changed its ID.')
    123
                if run['status'] in ACTIVE:
    124
                    time.sleep(15)
    125
            if run['status'] != 'SUCCEEDED':
    126
                raise RuntimeError('Actor ended with status ' + run['status'] + '.')
    127
            dataset_id = run.get('defaultDatasetId')
    128
            if not _id(dataset_id):
    129
                raise RuntimeError('Successful run has no valid dataset ID.')
    130
            rows = _request(opener, token, '/datasets/' + dataset_id + '/items?format=json&clean=true&limit=100')
    131
            if not isinstance(rows, list) or not rows:
    132
                raise RuntimeError('Dataset is empty or malformed.')
    133
            reviewed = []
    134
            covered = set()
    135
            for row in rows:
    136
                if not isinstance(row, dict):
    137
                    raise RuntimeError('Dataset contains a malformed row.')
    138
                try:
    139
                    domain = _domain(row.get('domain'))
    140
                except ValueError:
    141
                    raise RuntimeError('Dataset row has no valid domain.') from None
    142
                covered.add(domain)
    143
                reviewed.append({**row, 'needsReview': domain not in requested or row.get('confidence') != 'high'
    144
                                 or not _linkedin(row.get('linkedinUrl'))})
    145
            missing = [domain for domain in requested if domain not in covered]
    146
            return {'runId': run_id, 'datasetId': dataset_id, 'rows': reviewed,
    147
                    'missingDomains': missing,
    148
                    'needsReview': bool(missing) or any(row['needsReview'] for row in reviewed)}
    149
        except RuntimeError as error:
    150
            raise RuntimeError(f'Run {run_id}: {error}') from None
    151