How to read this
The dashboard scrapes nothing itself. It rents every fetch. Each tab is one platform and carries the same four sections, in the same order: Current flow (what runs today, with links) → Misconceptions (what is wired but never used, and where the recorded price is wrong) → Actors and data sets (the full table) → Flow diagrams (current and suggested, each carrying price and data limits).
Status is what the code does, which is not always what the registry says. apidojo/tweet-scraper
shows as live-and-hardcoded while xquik shows as replies-only, even though the registry names xquik
primary for both. That gap is the point of several tabs.
| Shape | What it is | Who bills | Count |
|---|---|---|---|
| Actor | A rented scraper program on Apify.
Send a URL, get rows, pay per row. Lives at apify.com/<slug> | Apify | 30 |
| Dataset | Bright Data's equivalent, addressed by an id like
gd_lkaxegm826bjpoo9m5 against one endpoint. No public page exists per dataset | Bright Data | 8 (4 wired) |
| Proxy | Bright Data Web Unlocker — a residential fetch with JS rendering, used when a plain request is blocked | Bright Data | 1 |
| Free / local | Public feeds, oEmbed, RSS, and two programs running on your own server | nobody | 10 |
Two account facts that change every price here: the Apify plan is BRONZE, measured rather than assumed, and tiered per-event pricing means the same actor costs a different amount on a different plan. Bright Data datasets bill per returned row with no per-run object to read a bill from, so their cost is rows × rate rather than a settled figure.
Every vendor, all platforms
What the status words mean
Status describes what the code actually does, which is not always what the registry says. Several vendors are registered and configured but can never be reached.
| Status | What it means |
|---|---|
| live | Runs in normal operation and we are billed for it |
| live — 302 of 310 | Runs, and the numbers say how much of the stored data it actually brought in. Where two vendors share a job, this shows which one is really doing the work |
| live, primary | First choice for this job, and it genuinely is the one that runs |
| live — hardcoded | Runs because its web address is typed directly into the code. The tier chooser is asked which actor to use and its answer is ignored, so the registry's choice has no effect |
| fallback | Second in line. Only runs if the first choice fails — in practice, rarely or never |
| unreachable | Registered as a backup but cannot ever run. Either the code only ever reads the first choice in the list, or the primary's address is hardcoded so no fallback is consulted. Listed, but decoration |
| disabled revert copy | A switched-off copy of an actor we stopped using, kept so we could go back to it if the replacement went badly. Deliberately inert — it cannot run and costs nothing |
| never wired | Available on the vendor's side and sometimes already format-checked, but no code calls it. Its real rate is unknown because it has never run |
| CLI only, never run | Reachable only by typing a command by hand. No scheduled job or dashboard button triggers it |
| free | A public feed, an official API within its free quota, or a program on our own server. No vendor, no bill |
How the prices work
Almost no vendor charges one flat price. A single run can combine four different charges, and our registry can only record the first of them — which is why the Extra charges & limits column below matters as much as the rate.
| Kind of charge | What it means | Worst example here |
|---|---|---|
| Per item | Charged for every row that comes back — one post, one comment, one follower, one ad. This is the only kind our registry can record | apify/facebook-pages-scraper at $0.0082 a page |
| Start fee | A flat charge every time the job runs, before any data arrives. Splitting one sweep into many small runs pays it many times over | data_xplorer/tiktok-ads-library-fast at $0.025 per run |
| Minimum | The smallest amount you are allowed to ask for. You pay for that many even when fewer exist — a quiet account still bills the full minimum | X followers: 200 people minimum · Threads: 10 posts |
| Overshoot | The vendor sends back more than you asked for and bills for all of it. Not a minimum — you asked for a small number and were charged for a larger one | One tweet billed as 15 rows (≈$0.0033, not $0.00022) |
| Conditional charge | An extra fee triggered by an option you turned on — a date filter, an email lookup, profile enrichment — or even by a run that finds nothing | LinkedIn no-result: $0.001 for an empty run |
The full list
| Vendor | Platform | Job | Status | Per item | Extra charges & limits |
|---|---|---|---|---|---|
| apify/instagram-scraper | posts, post detail, commenter profiles | live | $0.0023 | — | |
| apify/instagram-profile-scraper | profile + 12 posts free | live | $0.0023 | optional extra account detail +$0.006 | |
| scrapesmith/instagram-comments-scraper | comments | live | $0.0005 per comment | start fee $0.00005 | |
| scraping_solutions/instagram-scraper-followers-following-no-cookies | followers / following | live | $0.0007 per person | minimum 25 people | |
| apify/instagram-post-scraper | posts backup | unreachable | $0.0015 | primary's address is hardcoded, so this never runs | |
| datadoping/instagram-followers-scraper | followers | disabled revert copy | $0.0014 | rate is in the adapter, not the registry | |
| apify/instagram-comment-scraper | comments | disabled revert copy | $0.0023 | — | |
gd_lkaxegm826bjpoo9m5 | page posts | live — 7 of 310 | $0.0015 per post | 40.5 s — too slow to lead, so it rarely runs | |
| apify/facebook-posts-scraper | page posts | live — 302 of 310 | $0.004 per post | start fee $0.001 · +$0.001 when a date filter is used | |
| apify/facebook-pages-scraper | page + personal profile | live | $0.0082 per page | registry records $0.00246 — 3.3× under | |
| apify/facebook-comments-scraper | comments | live | $0.002 per comment | start fee $0.001 · +$0.001 when a filter is used | |
| danek/facebook-posts-fast | one post by URL | live | $0.00233 per post | start fee $0.00005 | |
| danek/facebook-posts-share-scraper | resharers | live | $0.00167 per person | start fee $0.00005 · overshoots — asked 10, billed 12 | |
| apify/facebook-ads-scraper | Meta ads | ad library | live | $0.005 per ad | — |
gd_lkay758p1eanlolqw8 | comments | never wired | — | never run, so the rate is unknown | |
| apidojo/tiktok-scraper | TikTok | posts + profile free, post detail | live | $0.0003 per post | cheapest post source anywhere here |
gd_l1villgoiiidt09ci | TikTok | profile + top posts | fallback | $0.0015 per profile | returns top posts, not newest — can miss a new post |
| apidojo/tiktok-comments-scraper | TikTok | comments | live | $0.0003 per comment | — |
| clockworks/tiktok-followers-scraper | TikTok | followers | live | $0.001 per person | start fee $0.001 — not in the registry · minimum 25 |
| clockworks/tiktok-comments-scraper | TikTok | comments | disabled revert copy | $0.001 | — |
| data_xplorer/tiktok-ads-library-fast | TikTok ads | ad library, EU/UK only | live | $0.0015 per ad | start fee $0.025 — the largest here · US is rejected |
| apidojo/tweet-scraper | X | posts | live — hardcoded | $0.0004 per tweet | runs instead of the registry's choice, at 2.67× the price |
| xquik/x-tweet-scraper | X | replies — and posts, per registry | live, replies only | $0.00015 per tweet | registry names it first choice for posts too; the code overrides that |
| kaitoeasyapi/twitter-x-data-tweet-scraper-pay-per-result-cheapest | X | one tweet | live | $0.00022 per row | overshoots badly — 1 tweet bills 15 rows ≈ $0.0033 |
| kaitoeasyapi/premium-x-follower-scraper-following-data | X | followers | live | $0.00015 per person | minimum 200 people — biggest line in the ledger, $3.77 of $9.55 |
| api-ninja/x-twitter-replies-retweets-scraper | X | retweeters | live | $0.00035 per person | start fee $0.01 · must set category: 'retweets' or it returns repliers |
| apidojo/twitter-replies-scraper | X | replies backup | unreachable | $0.0004 per item | +$0.014 per query · comment layer has one fixed actor per platform |
| seemuapps/x-tweet-retweeters-scraper | X | retweeters backup | unreachable | $0.0015 | code only ever reads the first choice |
| khadinakbar/x-tweet-engagement-scraper | X | replies + quotes + retweets | unreachable | $0.006 | third in a list that is never read past the first |
| seemuapps/x-quote-tweets-scraper | X | quote tweets | CLI only, never run | $0.003 | no scheduled job or button triggers it |
| data_direct/twitter-tweet-quotes-scraper | X | quotes backup | unreachable | $0.006 | — |
gd_lwxkxvnf1cynvib9co | X | one tweet by URL | never wired | — | cannot list an account's tweets, so it serves no capability |
| YouTube channel RSS | YouTube | newest videos | live — cap 15 | $0 | public feed holds only 15 videos — the cap that pushes first pulls to paid |
| YouTube Data API v3 | YouTube | likes, comment counts | live | $0 | free quota of 10,000 units a day · needs YOUTUBE_API_KEY |
| streamers/youtube-channel-scraper | YouTube | deep pull + channel stats | live — 197 of 207 | $0.001 per video | returns no like counts at all |
| streamers/youtube-comments-scraper | YouTube | comments | live | $0.0015 per comment | registry records $0.004 — 2.7× over |
gd_lk56epmy2i5g7lzu0k | YouTube | one video by URL | never wired | — | cannot list a channel's videos |
| harvestapi/linkedin-company-posts | company posts | live | $0.002 per post | start fee $0.00005 · $0.001 even when it finds nothing · floor of 8 · optional comment/reaction bundles $0.002 each | |
gd_l1vikfnt1wgvvqz95w | company profile | live, primary | $0.0015 per profile | cheaper than the Apify fallback, and genuinely runs first | |
| harvestapi/linkedin-company | company profile | fallback | $0.004 per profile | start fee $0.00005 | |
gd_l1viktl72bvl7bjuj0 | person profile | live, primary | $0.0015 per person | — | |
| harvestapi/linkedin-profile-scraper | person profile | fallback | $0.004 per person | $0.01 if an email is requested — 2.5× more | |
| harvestapi/linkedin-post-comments | comments | live | $0.002 per comment | registry records $0.004 — 2× over · optional profile enrichment $0.002–$0.01 | |
| apimaestro/linkedin-post-detail | one post | live | $0.005 flat per post | — | |
| datadoping/linkedin-post-repost-scraper | resharers | live | $0.0014 per person | minimum 10 — fewer is rejected outright | |
| apimaestro/linkedin-post-reshares | resharers backup | unreachable | $0.005 | second in a list read only to the first | |
| benjarapi/linkedin-company-posts | posts backup | unreachable | $0.002 | primary's address is hardcoded | |
| data_xplorer/linkedin-ad-library-scraper | LinkedIn ads | ad library | live | $0.001 per ad | registry records $0.0005 — 2× under |
gd_lyy3tktm25m4avu764 | one post by URL | never wired | — | cannot list an account's posts | |
| futurizerush/meta-threads-scraper | Threads | posts + full profile free | live — no fallback | $0.0025 per post | start fee $0.005 · minimum 10 posts · overshoots — a 10-post ask billed $0.045, which is 16 posts |
| automation-lab/google-ads-scraper | Google ads | ad library | live | $0.001 per ad | start fee $0.005 per run · the ad limit applies per advertiser, not per run |
| tesseract.js | Google ads | OCR the creative | live — only copy source | $0 | runs on our own server — no vendor, no key |
| Bright Data Web Unlocker | all | rescue fetch when blocked | live | $0.0025 per fetch | — |
Test spend, and the rates that have drifted
Every figure in this document came from a real run billed to the account. Total: $0.40.
| Platform | What was run | Data limit applied | Cost |
|---|---|---|---|
| profile, posts, apidojo candidate — twice, on two accounts | 12 posts per run | $0.0716 | |
| Bright Data posts, pages-scraper, Apify posts | 5 posts, 1 page | $0.037 | |
| TikTok | posts | 10 posts | $0.003 |
| X | apidojo vs xquik, same account | 10 tweets each | $0.0056 |
| YouTube | free RSS + API, paid actor | 12 videos | $0.012 |
| posts with bundle switches on | 10 posts, 5 comments, 5 reactions | $0.118 | |
| Threads | posts | 10 posts — the hard minimum | $0.045 |
| Google Ads | ads by advertiser id · ad-text comparison across 3 actors | 5–10 ads per run, 4 advertisers | $0.039 |
| Total | $0.40 | ||
The LinkedIn line is the expensive one because turning the bundle switches on turned a 10-post request into 59 billed rows. That is the finding, not an accident.
Seven rates the registry has wrong
The registry (scripts/lib/provider-tiers.js) is what the Capability Explorer quotes from and
what the budget guard reserves against. These seven no longer match the bills.
| Vendor | Registry | Actual | Direction |
|---|---|---|---|
apify/instagram-scraper | $0.004 | $0.0023 | over-quotes by 74% |
apify/instagram-post-scraper | $0.0023 | $0.0015 | over-quotes |
apify/facebook-pages-scraper | $0.00246 | $0.0082 | under-quotes 3.3× |
streamers/youtube-comments-scraper | $0.004 | $0.0015 | over-quotes 2.7× |
harvestapi/linkedin-post-comments | $0.004 | $0.002 | over-quotes 2× |
data_xplorer/linkedin-ad-library-scraper | $0.0005 | $0.001 | under-quotes 2× |
clockworks/tiktok-followers-scraper | $0.00102 | $0.001 + $0.001 start | start fee missing |
The registry has no way to record a start fee
provider-tiers.js can record a price per item and nothing else — there is no field for a flat
charge. But twelve actors charge one every time they run, and none of those fees has ever reached an estimate.
| Actor | Charge to start | Also charges |
|---|---|---|
data_xplorer/tiktok-ads-library-fast | $0.025 | — |
api-ninja/x-twitter-replies-retweets-scraper | $0.01 | — |
automation-lab/google-ads-scraper | $0.005 | — |
futurizerush/meta-threads-scraper | $0.005 | — |
apify/facebook-posts-scraper | $0.001 | filter-applied $0.001 |
apify/facebook-comments-scraper | $0.001 | filter-applied $0.001 |
clockworks/tiktok-followers-scraper | $0.001 | — |
harvestapi/linkedin-company-posts | $0.00005 | no-result $0.001 |
harvestapi/linkedin-company | $0.00005 | — |
scrapesmith/instagram-comments-scraper | $0.00005 | — |
danek/facebook-posts-fast | $0.00005 | — |
danek/facebook-posts-share-scraper | $0.00005 | — |
Two of these deserve attention beyond the missing field. filter-applied means
using a date filter costs money, and no-result means you are billed when a run comes back
empty.
The registry also disagrees with itself
apify/instagram-scraper is recorded at $0.004 on lines 85, 167 and 191, and at
$0.0023 on line 462 — in the same file. Which price the Explorer shows depends on which job you ask about. The
real price is $0.0023, so three of the four entries are wrong.
The four biggest wins, across all platforms
| Platform | Change | Effect |
|---|---|---|
| Read the 12 posts already inside the profile reply instead of buying them again | $0.0299 → $0.0023, 15.5 s → 4.2 s | |
| X | Delete the hardcoded actor URL and let the resolver pick xquik | 2.5× cheaper, 2× faster, one vendor instead of two |
| YouTube | Page the free RSS feed past 15 and enrich paid rows with the Data API too | free, 15× faster, fills likes which the paid actor leaves empty |
| Google Ads | Batch all 17 advertiser ids into one run | $0.005 once instead of $0.085 |
Current flow
Five jobs, four actors — apify/instagram-scraper is used twice for two different purposes.
A routine refresh makes two calls in sequence. First
apify/instagram-profile-scraper is asked
{"usernames":[handle]} and returns the profile — and 12 complete posts inside the same reply.
Then apify/instagram-scraper is asked for posts
and buys those same posts again. That is the whole story of this tab.
Comments come from scrapesmith/instagram-comments-scraper, which needs post URLs and therefore cannot run until posts exist. Commenter profiles then need usernames from the comments, so the chain is three deep: posts → comments → commenter profiles. Followers come from scraping_solutions/instagram-scraper-followers-following-no-cookies and need only the handle.
Bright Data serves nothing on Instagram except the Web Unlocker, which rescues a blocked single-post fetch.
Misconceptions
apify/instagram-scraper ·
$0.0276 → $0.0023 12× less · 15.5 s → 4.2 sapify/instagram-scraper ·
quoted $0.004 → real $0.0023 74% overfallbackOnly backup, because back then $0.0023 was cheaper than $0.004.pull-data.js line 38 writes the main actor's web address straight into the code. The chooser never gets
asked, so this backup has never run and cannot run.apify/instagram-post-scraper ·
$0.0015 unreachable · registry says $0.0023DISABLED_instagram_datadoping — a switched-off
copy kept so we could go back to it if the swap on 22 Aug 2026 went badly. It cannot run. Its $0.0014 price is written
in the adapter but not in the registry, which is the real (small) problem: if you ever wanted to switch back,
the Explorer could not price it.datadoping/instagram-followers-scraper ·
$0 — it never runsapidojo/instagram-scraper, is about
4.7× cheaper, so it looked like an easy saving.apidojo/instagram-scraper ·
rejected · would lose viewspull-commenter-profiles.js
was built. It calls the posts actor in details mode and records $0.004 per person looked up.apify/instagram-scraper →
apify/instagram-profile-scraper ·
$0.004 → $0.0023 43% lessActors and data sets
| Name, store title and status | Returns | Returned but not used | Wanted, but no actor returns it | Cost |
|---|---|---|---|---|
| apify/instagram-scraper
"Instagram Scraper"live — posts, post detail, commenter profiles Inputan account handle or post links, plus how many posts to fetch. A second mode takes a username to look up one person. ReturnsPost details — caption, likes, comment count, views, video link, hashtags, location and the owner, one row per post. |
idshortCodeurltypecaptiontimestamplikesCountcommentsCountvideoUrlvideoViewCountvideoPlayCountvideoDurationdisplayUrlimageschildPostshashtagsmentionstaggedUsersmusicInfoaudioUrlproductTypepaidPartnershipcoauthorProducerslocationIdlocationNamealtownerIdownerUsernameownerFullNamefirstCommentlatestCommentsisCommentsDisabledisPinneddimensionsWidth/HeightoriginalWidth/Height |
hashtags — tags and entity_tags tables are both 0 rows;
keywords are rebuilt from caption text instead · musicInfo — no column, and it is what makes a Reel travel ·
productType — no column, so a Reel cannot be told from a feed post ·
taggedUsersmentionsaudioUrlaltownerFullNameownerIdisCommentsDisabledisPinneddimensions*original*locationId — no columns ·
firstComment / latestComments — ignored because comments are bought separately |
shares, bookmarks_count, reactions, thread_id,
posting_tool, topic_tag — all 0/482. Instagram publishes none of them |
$0.0023/item registry says $0.004 |
| apify/instagram-profile-scraper
"Instagram Profile Scraper"live — profile Input {"usernames":["handle"]} — just the account name.ReturnsProfile details — followers, following, post count, bio, verified tick, business category, external links — and the 12 most recent posts in full. |
idusernamefullNamebiographyfollowersCountfollowsCountpostsCountverifiedprivateisBusinessAccountbusinessCategoryNamebusinessAddressexternalUrlexternalUrlsexternalUrlShimmedfbidprofilePicUrlprofilePicUrlHDhighlightReelCountigtvVideoCountlatestIgtvVideosjoinedRecentlyrelatedProfileslatestPosts[12] |
latestPosts[12] — twelve complete posts with video, views, hashtags, taggedUsers
and productType, discarded on every refresh · relatedProfiles — Instagram's own "similar
accounts", free competitor discovery ·
businessCategoryNamebusinessAddresshighlightReelCountigtvVideoCountlatestIgtvVideosjoinedRecentlyfbidprofilePicUrlHDexternalUrlsprivate |
contact_email, contact_phone, country_code — 0/10.
Not in this actor's output |
$0.0023/profile +$0.006 optional about-account |
| scrapesmith/instagram-comments-scraper
"Instagram Comments Scraper"live — comments Inputthe web address of a post, plus how many comments to fetch. ReturnsComment details — the text, when it was written, likes, reply count, nested replies, and who wrote it (name, verified, private, picture). |
commentIdtextusernameuserFullNameuserIdownerownerUsernameownerProfilePicUrltimestamp (unix)likesCountrepliesCountchildCommentCountrepliesisVerifiedisPrivateisMentionablehasLikedCommentmediaTypemediapostIdpostUrlcommentUrl |
replies — nested reply objects arrive free; parent_platform_comment_id
and threading_depth columns exist and are 0/123, so reply threads are bought and thrown away ·
hasLikedCommentisMentionablemediamediaType ·
owner / postId are duplicates of fields already kept |
account_type (0/123), shares, is_post_author — not in this actor's output |
$0.0005/comment +$0.00005 start |
| scraping_solutions/instagram-scraper-followers-following-no-cookies
"Instagram Followers/Following Scraper | No Login | No Cookies"live — followers and following Inputan account handle and a direction — followers or following. ReturnsA list of people — username, full name, private and verified flags, profile picture. Identity only: no bio, no follower counts. |
idusernamefull_nameis_privateis_verifiedprofile_pic_urltypeusername_scrape — 8 fields total |
type and username_scrape are echoes of your own request, not data |
The big gap. bio, followers_count, following_count,
posts_count, location, account_created_at are all 0/45 — the actor sends
none of them. Filling them needs a profile call per follower: about $58 for 25,000 |
$0.0007/person min 25 |
apify/instagram-post-scraper
"Instagram Post Scraper"unreachable — registered as a backup,
but pull-data.js hardcodes the primaryInputaccount handles or post links. ReturnsPost details, but without the video link — which is why it was kept as a backup only. |
posts, without videoUrl |
— (never runs) | — | $0.0015/post registry says $0.0023 |
datadoping/instagram-followers-scraper
"Instagram Followers Scraper (No Cookie)"disabled revert copy
— DISABLED_instagram_datadopingInputan account handle. ReturnsA follower list — 6 fields only: id, username, full name, private, verified, picture. Followers only — it cannot list who an account follows. |
idusernamefull_nameis_privateis_verifiedprofile_pic_url |
— (ran once, 45 rows) | followers only — no following direction at all |
$0.0014/person 2× the live actor; rate is in the adapter, not the registry |
apify/instagram-comment-scraper
"Instagram Comments Scraper"disabled revert copy
— DISABLED_instagram_apifyInputthe web address of a post. ReturnsComment details, but without the commenter's verified flag, private flag or profile picture. |
comments, without isVerified / isPrivate / profile picture |
— (never runs) | — | $0.0023/comment |
| TOTAL — measured test spend | 6 runs across 2 accounts, both rounds | Limits applied: 12 posts per run, 1 profile per run | $0.0716 |
Flow diagrams
current 4 steps, step 2 redundant
A routine incremental refresh of one Instagram account.
“Instagram Profile Scraper” — a rented scraper program. You give it a username, it hands back that account’s details — follower count, bio, verified tick — and its 12 most recent posts, at no extra charge.
{"usernames":["rollersoftware"]} → profile + latestPosts[12], complete with
videoUrl on 8/8 videos“Instagram Scraper” — a different rented program that fetches an account’s posts. It is the dearer of the two and returns posts we were already given free in step 1.
shortCode set to step 1. This is the entire overspend.“Instagram Comments Scraper” — you give it the web address of a post, it returns the comments under it with each commenter’s name. It cannot start until posts exist.
“Instagram Scraper” — the same posts program, run in a second mode that looks up a person instead. It charges $0.004 per person when the profile program does the same job for $0.0023.
pull-commenter-profiles.js calls instagram-scraper in details mode“Instagram Followers/Following Scraper | No Login | No Cookies” — give it a username and it lists the people who follow that account. It needs nothing from the other steps, so it can run on its own.
suggested 1 call, 2 rare branches
Read what already arrives; buy only the 3% the bundle cannot carry.
“Instagram Profile Scraper” — the same program as before — but this time we keep the 12 posts it already includes instead of throwing them away and buying them again.
“Instagram Scraper” — the dearer posts program, now used rarely rather than every time — just to fetch how long a video runs.
videoDuration → duration_seconds. Applies to 103/482 posts.“Instagram Scraper” — same program again, used only when a post is co-authored with another account, to fetch the co-author’s name. Affects 3 posts in 100.
coauthorProducers → coauthor_handles. Applies to 14/482 posts.“Instagram Comments Scraper” — unchanged, plus the followers program. The one change: commenter look-ups move to the cheaper profile program.
instagram-profile-scraper: same data, $0.0023 not $0.004.Field check — does the suggested flow lose anything?
Comparing the two live runs key by key, the bundle holds 14 of the 17 fields the dashboard stores. Every "how often" number below is counted straight from the database. The three gaps and their answers:
| Missing from bundle | Stored as | How often (DB-verified) | Answer |
|---|---|---|---|
paidPartnership | is_paid_partnership |
481/482 confirmed | No call needed — the bundle leaves the field out when the answer is "no". Treat missing as false |
videoDuration | duration_seconds |
103/482 confirmed | Top-up branch, only on a new video post. Skip the branch and you lose 103 rows |
coauthorProducers | coauthor_handles |
14/482 (3%) confirmed | Top-up branch, only on a new collab post. Skip the branch and you lose 14 rows |
So the honest claim is one call covers every stored field on about 97% of posts, with a rare second call for the rest — provided the two top-up branches are actually built. Leave them out and the saving is real but 117 rows of data stop being collected.
One thing to know about views: it is filled on only 102 of 482 Instagram posts, because
Instagram publishes a view count for videos only — and there are 103 videos. That does not change the decision to reject
apidojo, which returned no view count on any post, video or not.
Current flow
Seven jobs across two vendors — the only platform where Bright Data and Apify compete for the same work.
Posts have two tiers. The registry names Bright Data's gd_lkaxegm826bjpoo9m5 primary at $0.0015, with
apify/facebook-posts-scraper as the fallback.
In practice the fallback does almost everything — 302 of 310 stored posts — because Bright Data takes 40.5 s
against Apify's 5.3 s.
The page profile is a separate paid call to apify/facebook-pages-scraper: it is the only source of an exact follower count, and it returns no posts, so there is no Instagram-style bundle to exploit here.
Comments come from apify/facebook-comments-scraper and need post URLs from the posts job. One post by URL goes to danek/facebook-posts-fast, resharers to danek/facebook-posts-share-scraper, and the Meta ad library to apify/facebook-ads-scraper.
Misconceptions
gd_lkaxegm826bjpoo9m5 as the
first choice for Facebook posts, with the Apify actor as the backup.gd_lkaxegm826bjpoo9m5 (quoted) vs
apify/facebook-posts-scraper (actually used) ·
quoted $0.0015 → really $0.004 + $0.001 3.3× underapify/facebook-pages-scraper ·
quoted $0.00246 → real $0.0082 3.3× under-quotedcomments, not commentsCount, not topComments. That explains a number that
looked like a bug: comments_count is filled on only 69 of 310 Facebook posts. The 7 Bright Data
posts all have it; nearly all the Apify ones do not.apify/facebook-posts-scraper ·
no extra cost · the note is simply out of dateapify/facebook-pages-scraper ·
saving rejected · would lose exact follower countsvideo_view_count and count_reactions_type on every single row,
and nothing in our code reads them, so those columns stay empty on that path. The pages actor sends whether a page is
verified, and is_verified is empty for all 11 accounts. We already paid for all of it.apify/facebook-posts-scraper and apify/facebook-comments-scraper charge an extra
$0.001 every time a filter is applied, on top of a $0.001 charge just for starting. Our registry has no
place to record either kind of fee, so neither has ever appeared in an estimate.apify/facebook-posts-scraper ·
+$0.001 filter +$0.001 start never quotedgd_lkay758p1eanlolqw8 (Bright Data "Facebook Comments") holds full details about who commented, and
its format has already been checked against our database. No code calls it. So on the rare runs where Bright
Data fetches the posts, no comments come along with them.gd_lkay758p1eanlolqw8 · never run — rate unknownActors and data sets
| Name, store title and status | Returns | Returned but not used | Wanted, but no actor returns it | Cost |
|---|---|---|---|---|
gd_lkaxegm826bjpoo9m5Bright Data — "Facebook — Pages Posts by Profile URL"
served 7 of 310 declared primaryno public page; api.brightdata.com/datasets/v3/scrapeInputa Facebook page's web address and how many posts to fetch. ReturnsPost details — text, date, likes by type, comment and share counts, video views, attachments — plus 9–12 page profile fields free (followers, category, email, phone, website). |
44 keys. Post: post_idurlcontentdate_postednum_commentsnum_sharesnum_likes_typevideo_view_countcount_reactions_typeattachmentshashtagsis_sponsored
Page: page_followerspage_is_verifiedpage_emailpage_phonepage_categorypage_external_websitepage_namepage_intropage_creation_timepage_logopage_price_rangepage_reviews_score |
video_view_count and count_reactions_type — on every row,
never mapped, so views and reactions are 0/7 · all 9–12 page_* fields
— the page profile arrives free and is discarded |
Exact follower counts — page_followers is rounded (10,000 vs a real 10,895) |
$0.0015/post 40.5 s |
| apify/facebook-posts-scraper
"Facebook Posts Scraper"live — the de-facto primary, 302 of 310 Inputa Facebook page's web address and a post limit. ReturnsPost details — text, time, likes, shares, view counts, reactions, media, whether it is a video. No comment count at all. |
postIdtopLevelUrlurltexttimetimestamplikessharesviewsCountvideoPostViewCounttopReactionsCountmediaisVideopaidPartnershipcollaboratorspageNameuserfeedbackIdliveViewerCounttextReferencespageAdLibraryfacebookId |
collaboratorstextReferencesfeedbackIdliveViewerCountpageAdLibrary — no columns |
Comment counts. No comments, commentsCount or topComments in the
payload — comments_count is 61/302. The registry's "includes topComments" note is out of date |
$0.004/post + $0.001 start 5.3 s |
| apify/facebook-pages-scraper
"Facebook Pages Scraper"live — page and personal profiles Inputa Facebook page or personal profile address. ReturnsPage profile details — the exact follower count, likes, category, intro, email, website, messenger, creation date, cover photo. No posts. |
followers (10,895 exact)likesemailwebsitewebsitescategorycategoriesintroinfotitlepageIdpageNamepageUrlcreation_datemessengerprofilePictureUrlcoverPhotoUrlad_statusadditionalPropertiespageAdLibraryfacebookId |
messengercoverPhotoUrlad_statuscategoriesadditionalProperties
· verification — is_verified is 0/11 although it is returned |
Posts — no post array at all, so no bundle exists on Facebook | $0.0082/page registry says $0.00246 8.0 s |
| apify/facebook-comments-scraper
"Facebook Comments Scraper"live — the only Facebook comment source Inputthe web address of a single post. ReturnsComment details — text, date, likes, how deeply nested the reply is, and the commenter's name, profile link and picture. |
idtextdatelikesCountprofileNameprofileUrlprofilePicturecommentUrlthreadingDepthfacebookIdprofileId |
threadingDepthfacebookIdprofileId |
— | $0.002/comment + $0.001 start |
| danek/facebook-posts-fast
"Facebook Posts Fast"live — one post by URL Inputpost links passed as direct_urls — as objects, not plain text.ReturnsOne post in full — total reactions broken down by type, comments, reshares, and video with dimensions. |
total_reactions, per-type reaction breakdown, comments, reshares, video with dimensions |
per-type reaction split — stored only as a total | — | $0.00233/post input is direct_urls as objects, not strings |
| danek/facebook-posts-share-scraper
"Facebook Posts Share Scraper"live — resharers, Explorer panel only Inputthe web address of a post. ReturnsA list of people who reshared it — names with profile links. Fewer than the share count shows, because most Facebook shares are private. |
named resharer accounts with profile links | — | The true share count — Facebook shares are mostly private, so it returns fewer than posts.shares shows |
$0.00167 + $0.00005 start overshoots: asked 10, billed 12 |
| apify/facebook-ads-scraper
"Facebook Ads Library Scraper"live — Meta ad library Inputa page or search term for Meta's public ad library. ReturnsAd details — the full creative, body text, link, page id, ad format, page like count, call-to-action and display format. |
full creative, body text, linkUrl, pageId, ad format, page like count, cta_text, display_format |
ecommerce-enrichment events | — | $0.005/ad |
gd_lkay758p1eanlolqw8Bright Data — "Facebook — Comments"
never wiredInputthe web address of a post. Never actually called. ReturnsComment details with full commenter identity. Schema-checked against our database but no code uses it. |
full commenter identity | — (never runs) | — | — |
| TOTAL — measured test spend | 3 runs | Limits applied: 5 posts, 1 page | $0.037 |
Two things Facebook cannot do at any price: page followers are not enumerable (verified across three candidate actors — none returns the people), and a Page has no public following list. Both are platform limits.
Flow diagrams
current declared primary barely used
One account refresh. The tier order on paper is the reverse of the tier order in practice.
“Facebook — Pages Posts by Profile URL” — not an Apify program but a Bright Data dataset, called by its ID against one web address. There is no public page for it. Give it a page link, it returns that page’s posts.
page_* fields
free — all discarded, and page_followers is rounded.“Facebook Posts Scraper” — a rented program that reads a Facebook page’s posts. Eight times faster than the first choice, which is why it ends up doing nearly all the work.
comments_count is 61/302 on this path.“Facebook Pages Scraper” — returns a page’s own details — exact follower count, category, contact info. It returns no posts, so it cannot be merged with the step above.
“Facebook Comments Scraper” — give it a post’s web address and it returns the comments beneath it. The only Facebook comment source we have.
suggested same calls, correct order and prices
Nothing can be deleted here. What changes is which tier leads and what gets mapped.
“Facebook Posts Scraper” — the same program as today — the only change is that the price list stops pretending something else does this job.
“Facebook — Pages Posts by Profile URL” — stays as the spare in case the main one breaks. It already sends view counts and reaction breakdowns on every row that nothing reads.
video_view_count → views and count_reactions_type →
reactions. Both arrive on every row and neither is read.“Facebook Pages Scraper” — unchanged in what it does. It is the only source of an exact follower count, so it cannot be dropped — but it really costs $0.0082, not the $0.00246 on file.
is_verified, which is 0/11 despite being returned.“Facebook Comments Scraper” — no change. Optionally wire Bright Data’s unused comments dataset so comments also arrive on the Bright Data path.
gd_lkay758p1eanlolqw8 so Bright Data runs carry comments too.| Fix | Why | Cost |
|---|---|---|
| Make Apify the declared primary for posts | It already serves 302 of 310. Bright Data's 40.5 s is too slow to lead, and pretending otherwise prices every estimate against a tier that rarely runs | none |
Map video_view_count and count_reactions_type | Both arrive on every Bright Data
row and neither is read — views and reactions are 0/7 on that path | none |
Map is_verified from the pages actor | 0 of 11 accounts have it, though it is returned | none |
| Correct the pages rate to $0.0082 | Registry records $0.00246, and the job runs monthly so the error compounds quietly | none |
Current flow
Five jobs, and TikTok is already on the cheapest wiring in the project.
apidojo/tiktok-scraper does two jobs at once:
it takes a profile URL or a video URL in the same startUrls field, and it returns the full
channel object with every post — so the profile costs nothing extra. At $0.0003 a post it is the cheapest
post source of any platform here.
Bright Data's gd_l1villgoiiidt09ci sits behind it as a fallback. It works, but it returns top posts
ranked by engagement rather than a chronological history, so it can miss a brand-new post.
Comments come from apidojo/tiktok-comments-scraper and need post URLs. Followers come from clockworks/tiktok-followers-scraper and need only the handle. Ads come from data_xplorer/tiktok-ads-library-fast, searched by advertiser name rather than id.
Misconceptions
apidojo/tiktok-scraper ·
$0 extra — rides along with postsapidojo/tiktok-scraper ·
$0 — already paid for · on 100% of postsclockworks/tiktok-followers-scraper at
$0.00102 per person and no charge for starting.clockworks/tiktok-followers-scraper ·
quoted $0.00102 → real $0.001 + $0.001/run ~4% underreposts_count and quotes_count
columns for TikTok, which implies someone expected them to fill.US as a region you can pick.data_xplorer/tiktok-ads-library-fast ·
impossible · plus $0.025 to start any runActors and data sets
| Name, store title and status | Returns | Returned but not used | Wanted, but no actor returns it | Cost |
|---|---|---|---|---|
| apidojo/tiktok-scraper
"TikTok Scraper"live — account posts and one video by URL, same actor Input startUrls — an account link or a single video link, same field either way — plus a count.ReturnsPost details — views, likes, comments, shares, bookmarks, video, hashtags, the song (title, artist, album, cover) — and the whole account profile attached to every post. |
idtitleviewslikescommentssharesbookmarksvideouploadedAtuploadedAtFormattedhashtagssongchannelpoicollabInfosubtitleInformationpostPageimagesinputSource |
song — title, artist, album, duration, cover art, on every post, no column ·
hashtags — the tags and entity_tags tables are both 0 rows ·
poi (place), collabInfo, subtitleInformation, postPage — no columns ·
parts of channel beyond follower count |
reposts_count and quotes_count — 0/63. TikTok publishes neither.
thread_id, post_owner_handle, posting_tool also 0/63 |
$0.0003/post 6.0 s |
gd_l1villgoiiidt09ciBright Data — "TikTok — Profiles"
fallback, rarely runsno public page Inputan account's profile web address. ReturnsProfile details plus top posts — up to 17 best-performing and pinned videos with view counts and descriptions. Ranked by engagement, not newest-first. |
top_videos (17)top_posts_datapinned_posts, each with
video_urlplaycountdescriptionhashtags, plus profile metadata |
— (fallback only) | Chronological order. It returns top posts by engagement, so a brand-new post can be missed entirely | $0.0015/profile |
| apidojo/tiktok-comments-scraper
"TikTok Comments Scraper"live — comments Inputvideo web addresses. ReturnsComment details — text, date, likes, reply count, language, whether the creator liked it, and the commenter. |
idtextcreatedAtlikeCountreplyCountcommentLanguageuserisAuthorLiked |
commentLanguageisAuthorLiked — no columns |
— | $0.0003/comment 13.3× cheaper than the actor it replaced |
| clockworks/tiktok-followers-scraper
"TikTok Followers Scraper"live — followers Inputan account handle. ReturnsA follower list — username, nickname, bio, avatar, follower / following / post counts, verified flag and account age. Richer than the Instagram equivalent. |
username, nickname, bio, avatar, follower / following / post counts, verified flag, account age | — mapped more completely than Instagram's equivalent | — | $0.001/person + $0.001 start registry has $0.00102, no start fee |
| data_xplorer/tiktok-ads-library-fast
"TikTok Ads Library Fast"live — ad library Inputan advertiser name (not an ID) and a region. ReturnsAd details — creatives, advertiser and region. Europe and the UK only — a US region is rejected when the job runs. |
ad creatives, advertiser, region | — | US ads. Coverage is EU/EEA + UK only — the schema advertises a US region and the runtime
validator rejects it. Searches by advertiser name, not id |
$0.0015/ad + $0.025 start |
clockworks/tiktok-comments-scraper
"TikTok Comments Scraper"disabled revert copy
— DISABLED_tiktok_clockworksInputvideo web addresses. ReturnsComment details, but with a thinner commenter object than the live actor returns. |
comments, without the fuller user object | — (never runs) | — | $0.001/comment |
| TOTAL — measured test spend | 1 run | Limit applied: 10 posts | $0.003 |
Not obtainable at any price: per-post resharer lists. All three "TikTok Repost Scraper" actors take profile links and answer a different question.
Flow diagrams
current already on the cheapest wiring
One account refresh. Nothing here is mis-routed.
“TikTok Scraper” — a rented program that takes either an account link or a single video link in the same box. Every post it returns carries the whole account profile attached, so account stats cost nothing extra.
startUrls takes a profile URL or a video URL. The channel object rides along
on every post, so no separate profile call exists.“TikTok — Profiles” — a Bright Data dataset addressed by ID, with no public page. It returns an account’s best-performing posts rather than the newest ones.
“TikTok Comments Scraper” — give it a video’s web address, it returns the comments underneath. 13 times cheaper than the program it replaced.
“TikTok Followers Scraper” — give it a username, it lists the people following that account — with their bios, which the Instagram equivalent does not provide.
“TikTok Ads Library Fast” — searches TikTok’s public ad archive by advertiser name (not ID). The archive exists to satisfy a European transparency law, so it only covers Europe and the UK.
suggested no re-routing, two free captures
TikTok is the one platform already doing what Instagram should. Nothing moves vendor.
“TikTok Scraper” — nothing changes. Asking for a single post is enough to refresh account stats, because the profile rides along with every post.
channel is attached.“TikTok Scraper” — no new program and no new charge — the song title, artist, album and cover already arrive inside every post from the step above.
“TikTok Followers Scraper” — a bookkeeping fix, not a change to how it runs. Our price list says $0.00102 per person and no run charge; it really charges $0.001 per person plus $0.001 per run.
Current flow
Five jobs across four actors, and the best-mapped platform in the project — likes,
comments_count, shares, reposts_count, quotes_count,
bookmarks_count, thread_id and posting_tool are all 398/398.
Posts run through apidojo/tweet-scraper. The tier
resolver is called, but pull-data-x.js line 45 hardcodes the run URL, so the resolver's answer is
discarded — see the misconceptions below.
Replies are the one place xquik/x-tweet-scraper is
used, in conversationIds mode, and it reuses posts.thread_id so it needs the posts job first.
One tweet by id goes to kaitoeasyapi/twitter-x-data-tweet-scraper-pay-per-result-cheapest,
followers to kaitoeasyapi/premium-x-follower-scraper-following-data,
and retweeters to api-ninja/x-twitter-replies-retweets-scraper.
Bright Data cannot serve X at all: its gd_lwxkxvnf1cynvib9co dataset needs a specific tweet URL and cannot
enumerate a profile.
Misconceptions
pull-data-x.js line 45 has the dearer actor's
web address typed directly into it, so that is what runs — every time, no matter what the chooser says. I ran both
actors against the same account with ten tweets each and compared the real bills:apidojo/tweet-scraper →
xquik/x-tweet-scraper · $0.0004 →
$0.00015 2.67× less · 4.7 s → 2.5 s · fix is deleting one line| Cost for 10 tweets | Per tweet | Latency | |
|---|---|---|---|
apidojo/tweet-scraper — running today | $0.0040 | $0.00040 | 4.7 s |
xquik/x-tweet-scraper — registry primary | $0.0016 | $0.00015 | 2.5 s |
isQuote becomes isQuoteStatus, a two-line change. The other is real: the cheaper
actor does not send source, which is the app the tweet was posted from ("Twitter for iPhone").
That column is filled on all 398 posts — but no screen, chart or filter in the
dashboard reads it. If it stays unused, the switch costs nothing at all.posting_tool (398/398 filled, 0 readers) ·
keep apidojo as backup if it ever matters-1 — and charged for all fifteen. It ignores the setting that is supposed to limit how many results
come back. So the true price of one tweet is about $0.0033, roughly 15× the advertised rate.kaitoeasyapi/twitter-x-data-tweet-scraper-pay-per-result-cheapest ·
advertised $0.00022 → real ≈$0.0033 15× moreretweets, it defaults to replies — and hands back people who
replied instead of people who retweeted. It does not fail or warn; it just answers a different question. It also
charges $0.01 just to start, which our registry has no way to record.api-ninja/x-twitter-replies-retweets-scraper ·
$0.00035 + $0.01 start · must set category: 'retweets'kaitoeasyapi/premium-x-follower-scraper-following-data ·
$3.77 of $9.55 spent · minimum 200 per runActors and data sets
| Name, store title and status | Returns | Returned but not used | Wanted, but no actor returns it | Cost |
|---|---|---|---|---|
apidojo/tweet-scraper
"Tweet Scraper"live — hardcoded
in pull-data-x.js line 45, overriding the registryInputan account handle or search term, plus a tweet limit. ReturnsTweet details — text, date, likes, replies, retweets, quotes, bookmarks, views, media, and source (the app it was posted from). |
fullTexttextcreatedAtlikeCountreplyCountretweetCountquoteCountbookmarkCountviewCountconversationIdisRetweetisQuoteisPinnedisConversationControlledsourcemediaextendedEntitiesentitiescardplacelangpossiblySensitiveauthortwitterUrlretweetvideos |
cardplacelangpossiblySensitiveisConversationControlledextendedEntitiesentitiesauthor — no columns |
— | $0.0004/tweet 4.7 s |
| xquik/x-tweet-scraper
"X Tweet Scraper"live — replies only
the registry names it primary for posts too Inputan account handle for posts, or conversationIds to fetch the replies under a thread.ReturnsTweet details — all the same counts, plus reply-chain fields. Missing only source, the posting app. |
same counts, plus inReplyToIdinReplyToUserIdinReplyToUsernameisQuoteStatusisNoteTweetisTranslatabledisplayTextRangeeditviewStatepossiblySensitive |
isNoteTweetisTranslatabledisplayTextRangeeditviewState |
source — the posting tool. The one field apidojo has and this does not, and
posting_tool is filled 398/398 today |
$0.00015/tweet 2.5 s |
| kaitoeasyapi/twitter-x-data-tweet-scraper-pay-per-result-cheapest
"Twitter (X) Data Tweet Scraper Pay Per Result Cheapest"live — one tweet by id Inputa single tweet's ID. ReturnsOne tweet with live counts — but padded to 15 rows, 14 of them empty with the ID -1, and billed for all 15. |
the full tweet with live counts | — | A clean single result — it returns 15 rows for one tweet (14 filler rows with id -1) and bills for all 15 |
$0.00022 × 15 min ≈ $0.0033/post |
| kaitoeasyapi/premium-x-follower-scraper-following-data
"Premium X Follower Scraper — Following Data"live — followers and following Inputan account handle and a direction — followers or following. ReturnsA list of people — username, name, bio, follower count, verified flag, post count and account age. |
username, name, bio, followers_count, verified flag, post count, account age |
— richer than the Instagram follower actor | — | $0.00015/person min 200 $3.77 of $9.55 total ledger spend |
| api-ninja/x-twitter-replies-retweets-scraper
"X (Twitter) Replies & Retweets Scraper"live — retweeters, Explorer panel only Inputa tweet's web address and category, which must be set to retweets.ReturnsA list of people — screen name, name, follower count. Users only, never tweets. |
flat user objects: screen_namenamefollowers_count |
— | Reply tweets — it returns users, never tweets. category defaults to REPLIES and must be set to
retweets or it answers the wrong question |
$0.00035/item + $0.01 start |
| apidojo/twitter-replies-scraper
unreachable Inputa tweet's web address. ReturnsReply tweets. Never runs — the comment layer uses one fixed actor per platform. | replies | — | — | $0.0004 + $0.014/query |
| seemuapps/x-tweet-retweeters-scraper
unreachable (tier 2) Inputa tweet's web address. ReturnsA list of people who retweeted. Never runs — the code reads only the first choice. | retweeters | — | — | $0.0015 + $0.00005 |
| khadinakbar/x-tweet-engagement-scraper
unreachable (tier 3) Inputa tweet's web address. ReturnsReplies, quotes and retweets in one call — the only actor offering all three together. Never runs. | replies, quotes and retweets in one call | — | — | $0.006 |
| seemuapps/x-quote-tweets-scraper
CLI only, never run Inputa tweet's web address, typed by hand at a terminal. ReturnsQuote tweets. No scheduled job or dashboard button triggers it. | quote tweets | — | — | $0.003 |
| data_direct/twitter-tweet-quotes-scraper
unreachable (tier 2) Inputa tweet's web address. ReturnsQuote tweets. Registered as a second choice that is never reached. | quote tweets | — | — | $0.006 |
gd_lwxkxvnf1cynvib9coBright Data — "X — Posts by URL"
never wiredInputone tweet's exact web address. It cannot take an account handle. ReturnsOne tweet, enriched. Because it cannot list an account's tweets, Bright Data serves no X capability at all. | one tweet, enriched | — | Cannot enumerate a profile — it needs a specific status URL, which is why Bright Data serves no X capability | — |
| TOTAL — measured test spend | 2 runs, same account | Limit applied: 10 tweets each | $0.0056 |
Flow diagrams
current the resolver is consulted, then ignored
One account refresh. Two vendors doing one vendor's job.
xquik“Tweet Scraper” — a rented program that reads an account’s tweets. Its web address is typed directly into our code, so it runs no matter what the price list chooses.
source.“X Tweet Scraper” — the same kind of program, 2.67× cheaper and twice as fast. Today it is only trusted with replies, even though the price list names it first choice for posts as well.
conversationIds mode, reuses posts.thread_id, so it needs step 1 first.“Twitter (X) Data Tweet Scraper Pay Per Result Cheapest” — advertised as the cheapest per result. Ask it for one tweet and it returns 15 rows — 14 of them empty padding — and charges for all 15.
maxItems: 15 rows billed for 1 tweet.“Premium X Follower Scraper — Following Data” — lists who follows an account, with bios and follower counts. You cannot ask for fewer than 200 people, which is what makes it our biggest single bill.
suggested one vendor for posts and replies
The clearest fix in the project: delete one hardcoded line.
“Tweet Scraper” — its web address sits on line 45 of
pull-data-x.js. Deleting that line lets the price list make the choice it has already made.“X Tweet Scraper” — already proven on replies. It fills every column the dearer one does, except the app a tweet was posted from. Two field names differ and need renaming in two lines of code.
fullText → text,
isQuote → isQuoteStatus.“Tweet Scraper” — demoted to backup. Worth keeping only if the posting-app column ever starts being used — it is the single field the cheaper program does not send.
posting_tool starts mattering — source is the one field xquik lacks.“X (Twitter) Replies & Retweets Scraper” — it can return either repliers or retweeters. Left alone it quietly defaults to repliers — the wrong people — without any warning.
source, which nothing reads.posting_tool records "Twitter Web App" or "Twitter for iPhone". It is complete data, but no
tab, chart or filter reads it. If it stays unused, the switch costs nothing at all.
Current flow
YouTube is the only one of the seven networks with a genuinely free enumeration route, and it is wired free-first.
A pull starts by asking fetchFree(). That resolves the channel id once (cached on the account), reads the
public RSS feed at feeds/videos.xml?channel_id=…, then enriches each video through the official
YouTube Data API v3 to fill the like and comment
counts RSS omits. No key for RSS; the Data API needs YOUTUBE_API_KEY and gives 10,000 units a day free.
fetchFree() returns null the moment limit > 15 — the RSS feed carries only
15 entries. FIRST_PULL_LIMIT is 30, so every initial backfill falls straight through to
streamers/youtube-channel-scraper.
Incrementals ask for 8 and correctly stay free.
Comments are a separate paid call to
streamers/youtube-comments-scraper.
Bright Data cannot serve YouTube at all — gd_lk56epmy2i5g7lzu0k needs a specific video URL and cannot
enumerate a channel.
Misconceptions
| Cost | Latency | Views | Likes | Duration | |
|---|---|---|---|---|---|
| Free RSS + Data API | $0 | 0.7 s | 12/12 | 12/12 | 10/12 |
streamers/youtube-channel-scraper | $0.012 | 11.0 s | 12/12 | 0/12 | 12/12 |
15× faster, free, and the only route that returns like counts. The paid actor wins on exactly two things: more than 15 videos, and the channel stats block.
if (limit > 15) return null; — meaning "if you want more than 15, give up on free entirely". The first
time we pull an account we ask for 30. So every first pull jumps straight to paying, for all 30, instead of
taking 15 free and buying only the rest.streamers/youtube-channel-scraper · $0 →
$0.001/video · 0.7 s vs 11.0 slikes is empty on most YouTube posts, which looks like the data is
simply not available.likes is filled on 41 of 207 posts, which lines up exactly with how rarely the free route is used.streamers/youtube-comments-scraper ·
quoted $0.004 → real $0.0015 2.7× overActors and data sets
| Name, source and status | Returns | Returned but not used | Wanted, but no source returns it | Cost |
|---|---|---|---|---|
YouTube channel RSSyoutube.com/feeds/videos.xml?channel_id=…
live — tried first, freeInputa channel ID, appended to a public feed address. No key, no account. ReturnsThe 15 newest videos — id, link, title, publish date, view count, thumbnail. No likes, no comment counts, and never more than 15. |
videoIdurltitlepublishedviewsthumbnailUrl |
— | More than 15 videos. A hard feed ceiling, which is why fetchFree() bails above 15 and first
pulls go paid |
$0 · 0.7 s |
YouTube Data API v3
googleapis.com/youtube/v3live — enriches the row aboveInputthe video IDs taken from the feed above, plus a free API key. ReturnsThe numbers the feed leaves out — like count, comment count and duration. An add-on to the row above, not a separate source. |
likeCountcommentCountduration |
— | Nothing extra — it is an enrichment of the RSS row, not a separate route | $0, 10k units/day needs YOUTUBE_API_KEY |
| streamers/youtube-channel-scraper
"Fast YouTube Channel Scraper"live — 197 of 207 stored posts Inputa channel's web address and how many videos to fetch. ReturnsVideo details plus channel statistics — views, duration, title, date, and the channel's subscriber count, total views, join date and description. No like counts. |
viewCountdurationtitledateurlvideoIdnumberOfSubscriberschannelTotalViewschannelTotalVideosisChannelVerifiedchannelJoinedDatechannelLocationchannelDescriptionchannelDescriptionLinksaboutChannelInfochannelAvatarUrlchannelBannerUrlisAgeRestricted |
channelBannerUrlchannelDescriptionLinksisAgeRestrictedchannelLocation
— no columns · the channel block is otherwise the reason to keep this actor |
Like counts. Returned likes on 0 of 12 videos in a live run. likes is
41/207 overall |
$0.001/video 11.0 s |
| streamers/youtube-comments-scraper
"YouTube Comments Scraper"live — comments Inputa video's web address and a sort order — top comments or newest first. ReturnsComment details — id, text, author, date, likes and replies. |
comment id, text, author, date, likes, replies; sortCommentsBy TOP_COMMENTS or NEWEST_FIRST
(verified working) |
— | — | $0.0015/comment registry says $0.004 |
gd_lk56epmy2i5g7lzu0kBright Data — "YouTube — Videos"
never wiredInputone video's exact web address. It cannot take a channel. ReturnsOne video, enriched. Because it cannot list a channel's videos, Bright Data serves no YouTube capability. |
one video, enriched | — | Cannot enumerate a channel — it needs a specific video URL, which is why Bright Data serves no YouTube capability | — |
| TOTAL — measured test spend | 2 runs (1 free, 1 paid) | Limit applied: 12 videos each | $0.012 |
Not obtainable at any price: subscriber lists and a channel's subscriptions. Both are private by platform design — every candidate actor returns a count, not the people.
Flow diagrams
current free below 15 videos, paid above
The branch that decides everything is a single if.
“YouTube public feed” — not a rented program at all — a public feed YouTube publishes for every channel. No key, no bill, no account.
if (limit > 15) return null;“Google’s official API” — Google’s own official service, used to add the like and comment counts the public feed leaves out. Free up to 10,000 requests a day.
“Fast YouTube Channel Scraper” — a rented program that reads a channel’s full back-catalogue. It is the only way past the 15-video limit — but it returns no like counts at all.
“YouTube Comments Scraper” — give it a video link, it returns the comments, newest-first or top-first. Really costs $0.0015 each, not the $0.004 on file.
likes is 41/207 as a direct result.suggested free covers more, paid keeps the backfill
No vendor change. Use the free route further and enrich both paths.
“YouTube public feed” — take all 15 free videos first, then buy only the remainder — instead of giving up on free entirely and paying for all 30.
“Google’s official API” — the same free Google service already used on the free path. Running it on paid rows too costs nothing and fills a column that is empty on 166 of 207 posts.
“Fast YouTube Channel Scraper” — still needed for two things nothing else gives: more than 15 videos, and the channel’s own figures such as subscriber count.
numberOfSubscribers + channel block.“YouTube Comments Scraper” — bookkeeping only. The recorded $0.004 is 2.7× the real price, making YouTube comments look far dearer than they are.
likes stops being empty. Three fixes, none of which
costs anything.Current flow
Nine jobs across two vendors, and the most expensive platform per unit — LinkedIn is the hardest network to scrape, which is why competition is thin and prices are higher.
Company posts come from harvestapi/linkedin-company-posts, which has a floor of 8: asking for 2 returned 8 and billed for 8.
Both profile jobs lead with Bright Data, and that position is earned on price rather than assumed —
gd_l1vikfnt1wgvvqz95w for a company and gd_l1viktl72bvl7bjuj0 for a person, both at $0.0015
against Apify's $0.004. The harvestapi actors sit behind them as fallbacks.
Comments come from harvestapi/linkedin-post-comments and need post URLs, so the chain runs three deep: posts → comments → commenter profile. One post by URL goes to apimaestro/linkedin-post-detail, resharers to datadoping/linkedin-post-repost-scraper, and ads to data_xplorer/linkedin-ad-library-scraper.
Bright Data's LinkedIn Posts dataset gd_lyy3tktm25m4avu764 is never wired — it needs a specific
post URL and cannot discover an account's posts.
Misconceptions
scrapeComments / scrapeReactions as having
run and returned nothing. Apify now prices comment and reaction events on that actor, so both were
re-run with maxComments: 5 and maxReactions: 5 against a 10-post ask:| Items | Charged | Cost | |
|---|---|---|---|
| Posts only (10) | 10 | post: 10 | $0.0201 |
| Posts + bundles on | 59 | post: 10, reaction: 47, comment: 2 | $0.1181 |
5.9× the cost. Reactions bill at $0.002 each — the same as a post — and 47 arrived for a 10-post ask.
Worse, the extras come back as separate flat rows, not nested arrays: the first item's keys are
type, id, reactionType, actor, postId, query, a reaction record rather than a post. The existing parser would
need rewriting. And it yielded 2 comments across 10 posts, so it is a poor comment source too.
Leave both switches off — but rewrite the note, because "ignored" is wrong.
harvestapi/linkedin-company-posts · $0.0201 →
$0.1181 5.9× more · 10 posts asked → 59 rows billedharvestapi/linkedin-post-comments ·
quoted $0.004 → real $0.002 2× overdata_xplorer/linkedin-ad-library-scraper ·
quoted $0.0005 → real $0.001 2× under-quotedharvestapi/linkedin-company-posts has a charge called no-result at $0.001 — you pay
it when the run comes back empty. It also charges a small fee just to start. Our registry can record neither,
so a run that finds nothing currently looks free in every estimate.harvestapi/linkedin-company-posts · $0.001 for nothing never quotedField input.max_reposts must be >= 10). Someone already found these and wrote them into the code.
They are worth keeping.Actors and data sets
| Name, store title and status | Returns | Returned but not used | Wanted, but no actor returns it | Cost |
|---|---|---|---|---|
| harvestapi/linkedin-company-posts
"LinkedIn Company Posts Scraper"live — company posts Inputa company page's web address and a post limit. Optional switches can also pull comments and reactions. ReturnsPost details — content, date, engagement counts, author, media, images, video, article, repost info. Never fewer than about 8 posts. |
contentpostedAtengagementsocialContentstatsauthormediapostImagespostVideoarticlerepostrepostIdreshared_postis_resharedshareUrnshareLinkedinUrllinkedinUrlentityIdplatformPostIdcommentIdsreactionIdsnewsletterTitlenewsletterUrlcontentAttributes |
articlenewsletterTitlenewsletterUrlcontentAttributesshareUrnentityIdcommentIdsreactionIds — no columns |
View counts. views is 0/93 — LinkedIn publishes impressions only to the post owner |
$0.002/post + $0.00005 start floor of 8 bundles: reaction $0.002, comment $0.002 |
gd_l1vikfnt1wgvvqz95wBright Data — "LinkedIn — Company Profile"
live, primary — cheaper than Apifyno public page Inputa company page's web address. ReturnsCompany profile details — follower count, employee count, about text, specialities, industries, headquarters and founding year. |
followers, employees, about, specialties, industries, headquarters, founded year | — | — | $0.0015/profile |
| harvestapi/linkedin-company
"LinkedIn Company Scraper"fallback for the row above Inputa company page's web address. ReturnsCompany profile details — the same shape as the Bright Data row above, at 2.7× the price. Backup only. |
same shape | — (fallback only) | — | $0.004 + $0.00005 start |
gd_l1viktl72bvl7bjuj0Bright Data — "LinkedIn — Person Profile"
live, primaryno public page Inputan individual's profile web address. ReturnsPerson profile details — name, city, current position, about text, current company, work history and education. |
name, city, position, about, current_company, experience, education |
— | — | $0.0015/person |
| harvestapi/linkedin-profile-scraper
"LinkedIn Profile Scraper"fallback for the row above Inputan individual's profile web address, optionally asking for an email too. ReturnsPerson profile details, same shape as above, plus an email address if requested — which costs 2.5× more. |
same, plus optional email | — (fallback only) | — | $0.004 $0.01 with email |
| harvestapi/linkedin-post-comments
"LinkedIn Post Comments Scraper"live — comments Inputa post's web address. ReturnsComment details — id, text, author, date, likes and the commenter's job title, which no other network gives us. |
comment id, text, author, job title, date, likes — the job title is something no other network returns | main-profile and full-profile enrichment events | — | $0.002/comment registry says $0.004 |
| apimaestro/linkedin-post-detail
"LinkedIn Post Detail"live — one post by URL Inputone post's web address. ReturnsOne post in full — reactions broken down by type, comments, shares and video. |
reactions with a per-type breakdown, comments, shares, video | the per-type reaction split | — | $0.005 flat |
| datadoping/linkedin-post-repost-scraper
"LinkedIn Post Repost Scraper"live — resharers, Explorer panel only Inputa post's web address and a count that must be 10 or more — fewer is refused outright. ReturnsA list of people who reshared it — names, profile links and any commentary they added. |
named resharers with profile URLs and any commentary they added | — | — | $0.0014 hard floor of 10 — fewer is an HTTP 400 |
| data_xplorer/linkedin-ad-library-scraper
"LinkedIn Ad Library Scraper"live — ads (337 rows) Inputan advertiser to look up in LinkedIn's public ad library. ReturnsAd details — creatives and advertiser. 337 rows collected so far. |
ad creatives, advertiser | — | — | $0.001/ad registry says $0.0005 — under-quoted |
| apimaestro/linkedin-post-reshares
unreachable (tier 2) Inputa post's web address. ReturnsA list of people who reshared it — same shape as the live actor. Second in a list only ever read to the first. | same reposter shape | — | — | $0.005 |
| benjarapi/linkedin-company-posts
unreachable — the primary's URL is hardcoded Inputa company page's web address. ReturnsCompany post details. Never runs — the first choice's address is hardcoded. | company posts | — | — | $0.002 |
gd_lyy3tktm25m4avu764Bright Data — "LinkedIn — Posts"
never wiredInputone post's exact web address, matching /(pulse|posts|feed/update)/.ReturnsOne post, enriched. It cannot list an account's posts, so it is never wired in. | one post, enriched | — | Cannot enumerate an account — it needs a post URL matching /(pulse|posts|feed\/update)/ | — |
| TOTAL — measured test spend | 1 run, bundles on | Limits applied: 10 posts, 5 comments, 5 reactions → 59 rows billed | $0.118 |
Not offered: LinkedIn followers and following. A company-followers actor exists but has about 315 users and has never been run, so its rate is unknown and it is not wired.
Flow diagrams
current 5 calls, two vendors, three deep
One company refresh. The wiring is sound; the recorded prices are not.
“LinkedIn Company Posts Scraper” — give it a company page, it returns that company’s posts. It never returns fewer than about 8 — ask for 2 and you get, and pay for, 8.
“LinkedIn — Company Profile” — a Bright Data dataset called by ID, with no public page. Returns followers, employee count, industry and headquarters. Cheaper than the Apify equivalent, which is why it leads.
“LinkedIn Post Comments Scraper” — returns the comments under a post — including each commenter’s job title, which no other network gives us.
“LinkedIn — Person Profile” — a Bright Data dataset for an individual person — job, employer, history. Third link in the chain, so it cannot run until comments exist.
“LinkedIn Post Repost Scraper” — lists who reshared a post. It refuses any request for fewer than 10. Ads come from a separate program,
data_xplorer/linkedin-ad-library-scraper.suggested same wiring, corrected prices
LinkedIn consolidates into nothing. Every fix here is bookkeeping.
“no consolidation available” — no vendor bundles posts together with the company profile, and Bright Data already leads both profile jobs on price. There is no Instagram-style saving to find here.
“LinkedIn Post Comments Scraper” — recorded at $0.004, really $0.002. And the ad program
data_xplorer/linkedin-ad-library-scraper is recorded at $0.0005 but really costs $0.001 — an under-estimate, the dangerous direction.“LinkedIn Company Posts Scraper” — it can optionally fetch reactions and comments alongside posts. Switching that on turned a 10-post request into 59 billed rows and cost 5.9× more.
“LinkedIn Post Repost Scraper” — these minimums are the vendors’ rules, not ours. The resharer program returns an outright error below 10; the posts program silently rounds up to 8.
Current flow
One call, one actor, no fallback. Threads is the project's only documented single point of failure — and its payload is the richest of any platform, about 80 keys.
futurizerush/meta-threads-scraper is
asked {"usernames":["nasa"], "max_posts": 10} and returns posts and the entire profile attached to every
post: followers_count, bio, is_verified, display_name,
profile_pic_url, user_id, plus emails, phones, bio_links
and external_links. All present on 10 of 10 posts.
max_posts has a hard minimum of 10, so every Threads pull costs $0.045 whether one new post exists
or ten.
Everything else on Threads is a dead end — comments, resharers and followers are all platform or actor limits rather than wiring gaps. Bright Data serves nothing here.
Misconceptions
futurizerush/meta-threads-scraper ·
expected $0.030 → billed $0.045 ~50% overfuturizerush/meta-threads-scraper ·
$0 extra today · a separate call would cost a full run"Making the seemingly impossible, possible. ✨" — and the
accounts.bio column is empty. It is exactly the same pattern as the Instagram posts we were buying
twice: the data arrives, already paid for, and gets dropped on the floor.futurizerush/meta-threads-scraper ·
$0 — already paid for · arrives on 100% of postsfuturizerush/meta-threads-scraper ·
~$67/month at 50 accounts → weekly cadence
up to 6× lesssleek_waveform/threads-scraper is dramatically cheaper, which
looks like an obvious switch.logical_scrapers/threads-post-scraper: it returned zero items and cannot list an account's posts.futurizerush/meta-threads-scraper ·
no such modeActors and data sets
| Name, store title and status | Returns | Returned but not used | Wanted, but no actor returns it | Cost |
|---|---|---|---|---|
| futurizerush/meta-threads-scraper
"Meta Threads Scraper"live — no fallback exists
posts and the whole profile Input {"usernames":["nasa"], "max_posts": 10} — an account name and a count that cannot go below 10.ReturnsPost details plus the entire profile, attached to every post — around 80 fields. Engagement (likes, replies, reposts, quotes, views), media, 20+ true/false flags, and the account's followers, bio, verified tick, email and phone. |
~80 keys. Engagement: like_countreply_countrepost_countquote_countshare_countview_countview_count_status — all 10/10Profile, on every post: followers_countbiois_verifieddisplay_nameprofile_pic_urlprofile_pic_hd_urluser_idusernameemailsphonesbio_linksexternal_linksprofile_urlprofile_tagsPost: text_contentpost_codepost_urlcreated_attopic_taghashtagsmentionsmentioned_accountslanguagemedia_urlmedia_urlsmedia_typemedia_width/heighthas_mediahas_audioaccessibility_captionis_ai_generatedis_editedis_gifis_paid_partnershipis_pinnedis_spoileris_sticker_postis_quote_postis_replyis_repostreply_controllocation_idlocation_namepodcast_namepodcast_urlpodcast_platformsticker_idssticker_urlsurlsplus 7 replied_to_* and 4 reposted_* fields |
bio — on every post, accounts.bio is empty ·
emailsphonesbio_linksexternal_links — returned as arrays, empty
for NASA, and nothing maps them ·
accessibility_captionlanguageis_ai_generatedis_editedis_spoileris_sticker_postreply_controlmedia_width/heightpodcast_*sticker_*profile_tags
· the replied_to_* and reposted_* blocks |
Nothing extra from this actor — the gaps are the three platform limits below | $0.0025/post + $0.005 to start minimum 10 posts billed $0.045 for a 10-post ask 8.9 s |
logical_scrapers/threads-post-scrapertested and rejectedInputone post's web address. ReturnsNothing — it returned zero items on a real run, and it cannot list an account's posts at all. |
zero items on a real run | — | Cannot enumerate an account — it scrapes one post by URL | — |
sleek_waveform/threads-scrapertested and rejectedInputan account name. ReturnsPost details, 450× cheaper — but only ~18 people used it in 30 days, so nobody maintains it. |
posts | — | Maintenance — 18 users in 30 days. 450× cheaper and not dependable for the project's most fragile capability | — |
| TOTAL — measured test spend | 1 run | Limit applied: 10 posts — the hard minimum, so this is the cheapest possible Threads run | $0.045 |
Three things Threads cannot do, all verified rather than assumed
| Capability | Why not |
|---|---|
| Comments | The actor has User Posts / Reposts / Replies / Search modes — none is "replies TO a given post" |
| Who reposted | The only candidate takes profiles and returns what an account has
reposted — the opposite direction. Counts are stored free |
| Followers | Candidate actors exist with ~8 users in 30 days. Not offered, deliberately |
Flow diagrams
current one call, three dead ends, no fallback
The whole platform is a single actor.
“Meta Threads Scraper” — the only program we have for Threads — there is no spare. Give it a username and every post comes back with the account’s full profile attached, around 80 fields in total.
{"usernames":["nasa"], "max_posts": 10} → posts + full profile on every post.
~80 keys, the richest payload in the project.Nothing can be bought for this. The Threads program offers four modes — an account’s posts, its reposts, its replies, and search — and none of them fetches the replies underneath a given post.
The only program offering this takes an account and tells you what that account has reshared — the opposite of what we need. The number of reshares does arrive free with each post.
Programs claiming to do this have roughly 8 users in 30 days, meaning nobody maintains them. Deliberately not used, since Threads already has no backup.
suggested map what arrives, pace against the floor
There is nowhere to route to. The fix is what you keep, and how often you ask.
“Meta Threads Scraper” — no new program, no new charge. The bio already arrives attached to all 10 of 10 posts and is being dropped.
“Meta Threads Scraper” — because it refuses to return fewer than 10 posts, any extra call — even just for profile details — triggers the full 10-post charge again.
“Meta Threads Scraper” — every run bills for 10 posts whether or not 10 exist. For a quiet account, checking daily pays ten times over for nothing.
Current flow
One actor plus a local OCR step. 427 Google ads stored.
automation-lab/google-ads-scraper is
asked for ads from the Google Ads Transparency Center. Where an advertiser id is on file it uses
{advertiserIds: ["AR137…"], maxAds: 10} — an exact lookup; otherwise it falls back to
{searchTerms: [name]}, a fuzzy name search. There are 17 ids on file in
google_advertiser_ids.
What comes back is a picture, not words. So every creative image is then read by
tesseract.js running locally — no vendor, no key,
no bill — using the Worker API with { blocks: true } to get line-level positions rather than a flat text blob.
The result reaches ads with ad_copy_source set to 'ocr'. Bright Data has no
ad-library endpoint for any network, so all four ad capabilities in the project are legitimately single-tier.
null every time, so today the only wording we hold comes from
reading pictures. A different actor returns that wording properly on the same ads.s-r/google-ads-transparency alongside the current one ·
OCR text, 70/76 → clean title + body + display URLMisconceptions
adText, headline, body and callToAction and
returns all four empty — even on an ad Google itself labels format Text. The obvious conclusion was that
Google publishes ads as finished pictures and the wording simply is not there to fetch.Text, with 0 of 5 carrying any wording.
s-r/google-ads-transparency returned the
wording properly:title → "Entire House / Apartment in Switzerland, Europe — from $31 | trivago"ad_text → "Your Ideal Hotel in Switzerland. Compare Prices and Save on your Next Stay!"display_url → "http://www.trivago.com"Repeated on a second advertiser with the same outcome. Google does publish this. Our actor does not read it.
automation-lab/google-ads-scraper 0 of 5 vs
s-r/google-ads-transparency returns title + body + display URL| Format | In our database | Has wording today | Is real wording available? |
|---|---|---|---|
| text | 76 | 70 — via OCR | Yes. Google publishes it; our actor misses it, s-r returns it |
| image | 333 | 96 — via OCR | No. Genuinely a picture — OCR is the only route |
| video | 18 | 1 — via OCR | No. Same as image |
maxAds: 10 reads like "this run will fetch at most 10 ads".maxAds: 30 — it came
back with 300 items. So a brand with many IDs costs IDs × maxAds × price, and the real total can be
far larger than the number you typed.automation-lab/google-ads-scraper ·
10 IDs × 30 = 300 items · not 30automation-lab/google-ads-scraper ·
$0.085 in fees → $0.005 saves $0.08 a sweepAR12910272019299303425 — not one
of the six Expedia IDs we track. Big brands run a separate ad account per country, so the two lookups return
different ads. And even when it works, ad_text was filled on only 1 of 5 rows.s-r/google-ads-transparency ·
rejected as a replacement · accepted as a text-ad second passsolidcode/ads-transparency-scraper is the most-used Google ads
actor on the store, so it looked like the natural upgrade.solidcode/ads-transparency-scraper ·
rejected twicecta_text column being empty for Google looks like a collection
gap on our side.cta_text and display_format are empty on
all 427 Google ads — while Meta fills them on 97 and 104 of 115. Same columns, same code, different
platform. Google simply does not publish them.Actors and data sets
| Name, store title and status | Returns | Returned but not used | Wanted, but no source returns it | Cost |
|---|---|---|---|---|
| automation-lab/google-ads-scraper
"Google Ads Scraper"live — the only Google ad source Input advertiserIds (we hold 17) or searchTerms, plus maxAds — which applies per advertiser, not per run.ReturnsAd details — advertiser, creative id, format, first and last shown dates, image, media links, YouTube link. The four wording fields are returned empty every time. |
advertiserIdadvertiserNameisVerifiedcreativeIdadFormatfirstShownlastShownimageUrlmediaUrlspreviewUrlregionyoutubeIdyoutubeWatchUrlvideoRecoveredBy
· declared but always null: adTextheadlinebodycallToActiondestinationUrl |
isVerifiedcreativeIdregionpreviewUrl — no columns |
Ad copy — the actor's fault, not Google's. adText, headline, body,
callToAction are declared and always null, including on ads whose adFormat is
"Text", where Google does publish wording and s-r returns it ·
cta_text 0/427 and display_format 0/427 — these two Google really does not publish |
$0.001/ad + $0.005 per RUN maxAds is per advertiser id22.4 s |
| s-r/google-ads-transparency
"Google Ads Transparency Scraper - Ad Library & Creatives"
not currently used — proposed second pass for text ads only Inputa website domain (advertiser IDs and brand-name search both return nothing), optionally filtered to format: text.ReturnsAd details including the real wording — title, ad_text and display_url for text ads, plus advertiser, format, dates and image. |
advertiser_idcreative_idadvertiser_nameformatfirst_shownlast_showndays_showntarget_domaindeeplink
· the wording ours misses: titlead_textdescriptiondisplay_urlkeyword
· image_urlpreview_urlad_group_idcustomer_id |
— not wired yet | Advertiser IDs. advertiser_id returned 0 ads on two different IDs, and brand-name
search returned 0 — only a domain lookup works, and we store 17 IDs and no domains. Its domain lookup
also found a different Expedia account than the ones we track · ad_text filled on only 1 of 5 rows |
pay per result $0.0076 for 5 text ads 21.5 s |
| tesseract.js
local OCR, not a vendorlive — only wording source today Inputa creative image file, already downloaded. ReturnsThe words read out of the picture — line by line, with positions and a confidence score. Succeeds on about 39% of ads. |
line-level text with positions and a confidence score, via the Worker API with { blocks: true } |
— | The 351 image and video ads. OCR recovers copy on 167 of 427 — about 39% — and for images and video it stays the only route. On the 76 text ads it already gets 70, but as OCR'd picture text rather than clean fields | $0 · local, no key, no API |
solidcode/ads-transparency-scrapertrialled and rejected, v1.51Inputsearch terms only — it has no advertiser-ID input at all. ReturnsAd details, but 0 of 8 ads carried any video information. |
ads | — | 0 of 8 ads carried any video field, and it has no advertiserId input at all — recorded so it is
not re-evaluated a third time |
— |
| Bright Data Input— nothing to send. Bright Data has no ad-library endpoint for any network.
Returns— nothing. This is why all four ad capabilities in the project have a single supplier and no backup.
| — | — | No ad-library endpoint for any network. This is why all four ad capabilities in the project are single-tier | — |
| TOTAL — measured test spend | 11 runs | Limits applied: 5–10 ads per run, across 4 advertisers, comparing our actor against two alternatives on the same text ads | $0.039 $0.009 first round + $0.030 ad-text comparison |
What actually reaches the database, across all three ad networks
| Network | Ads | ad_copy_text | Landing url | Image | Video | cta_text |
|---|---|---|---|---|---|---|
| 427 | 167 — all from OCR | 427 | 415 | 18 | 0 | |
| meta | 115 | 113 | 100 | 115 | 58 | 97 |
| 18 | 18 | 18 | 18 | 9 | 0 |
Flow diagrams
current actor gives pictures, OCR gives words
One advertiser sweep.
advertiserIds where an id is on file (17 are), otherwise searchTerms — a
fuzzy name search that can drift onto a similarly-named brand.“Google Ads Scraper” — a rented program that reads Google’s public Ads Transparency Center. The only one that accepts an advertiser ID, which is how our whole pipeline is keyed.
adText, headline, body, callToAction all null“open-source text recognition” — not a vendor and not a rented program — a text-recognition library running on our own machine. No key, no account, no bill.
{ blocks: true } for line positions. Writes ad_copy_source = 'ocr'.
Doing picture-reading even on the 76 text ads, whose words are available as clean fields.suggested batch the fee, add a text pass, keep OCR
Keep our actor as the spine — it is the only one that accepts advertiser IDs.
“Google Ads Scraper” — it charges $0.005 each time it starts, but its ad limit applies per advertiser. So one run covering 17 advertisers pays that charge once instead of seventeen times.
maxAds is per advertiser id — so one run covers all
17 and pays the fee once.“Google Ads Scraper” — an exact ID cannot drift onto a similarly-named brand, and an ID lookup respects the ad limit where a name search does not.
maxAds
where a name search is not.“Google Ads Transparency Scraper - Ad Library & Creatives” — a different rented program that does return the real wording of text ads — headline, body and display URL as separate fields. It can only be looked up by website domain, never by advertiser ID.
format: 'text' for the advertisers that run search ads.
It returns title, ad_text and display_url as clean fields — wording we
currently guess from a picture, plus a display URL we store nowhere.“open-source text recognition” — for image and video ads there is no published wording anywhere, so reading the picture remains the only route and stays correct.