instagram-scraper
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
ChineseInstagram Scraper
Instagram Scraper
Public Instagram data via Apify's Instagram Scraper. No Instagram login, no cookies, no Meta app review.
| Actor | |
| Auth | |
| Price | from $2.30 per 1,000 results, billed by Apify |
| Free tier | $5/month credits ≈ 2,000+ results, no credit card |
通过Apify的Instagram Scraper获取公开的Instagram数据。无需登录Instagram,无需Cookie,无需Meta应用审核。
| Actor | |
| 身份验证 | |
| 价格 | 每1000条结果起价2.30美元,由Apify计费 |
| 免费额度 | 每月5美元额度 ≈ 2000+条结果,无需信用卡 |
Setup — do this before any data call
设置——在任何数据调用前完成此步骤
Always run this preflight first. Do not attempt a data call until it passes.
bash
if [ -n "$APIFY_TOKEN" ]; then
echo "AUTH_OK curl"
elif command -v apify >/dev/null 2>&1 && apify info >/dev/null 2>&1; then
echo "AUTH_OK cli: apify"
elif { [ -f "$HOME/.apify/auth.json" ] || [ -f "$USERPROFILE/.apify/auth.json" ]; } \
&& npx --yes apify-cli@latest info >/dev/null 2>&1; then
echo "AUTH_OK cli: npx --yes apify-cli@latest"
else
echo "AUTH_MISSING"
fiThe checks are ordered cheapest-first. The fallback costs about four seconds,
so it only runs when an Apify config directory shows a previous login — a first-time
user reaches instantly rather than waiting for a package download that
was never going to find a session.
npxAUTH_MISSING$HOME$USERPROFILE$HOME$USERPROFILE\.apify$HOMEAUTH_MISSINGAUTH_MISSINGThe two modes are not interchangeable — the preflight tells you which
call form to use.
AUTH_OK- — a token is in the environment. The HTTP calls in this skill work as written.
AUTH_OK curl - — the Apify CLI holds the session and the token is not readable from disk.
AUTH_OK cli: <prefix>carries account metadata only (username, plan, proxy groups — no~/.apify/auth.jsonfield); current CLI versions keep the token in the OS secrets backend. Do not try to extract one from that file: a bogustokenheader returnsAuthorization: Bearerand looks exactly like a revoked token. Use401instead, prefixed with whatever the preflight printed afterapify call(cli:, orapifywhen the CLI is not onnpx --yes apify-cli@latest).PATH
请始终先运行此预检步骤。在预检通过前,请勿尝试进行数据调用。
bash
if [ -n "$APIFY_TOKEN" ]; then
echo "AUTH_OK curl"
elif command -v apify >/dev/null 2>&1 && apify info >/dev/null 2>&1; then
echo "AUTH_OK cli: apify"
elif { [ -f "$HOME/.apify/auth.json" ] || [ -f "$USERPROFILE/.apify/auth.json" ]; } \
&& npx --yes apify-cli@latest info >/dev/null 2>&1; then
echo "AUTH_OK cli: npx --yes apify-cli@latest"
else
echo "AUTH_MISSING"
fi检查顺序按成本从低到高排列。回退方案耗时约4秒,因此仅当Apify配置目录显示存在之前的登录记录时才会运行——首次使用的用户会直接得到结果,无需等待永远找不到会话的包下载过程。
npxAUTH_MISSING**特意同时检查和。**在Windows系统中,它们可能指向不同位置——沙箱和部分CI镜像会重新映射,但Apify CLI仍会将内容写入。如果仅测试,可能会导致已正常登录的用户收到提示,进而引导他们注册已有的账户。如果在确认已认证的机器上看到,请先检查这两个路径再判断结果。
$HOME$USERPROFILE$HOME$USERPROFILE\.apify$HOMEAUTH_MISSINGAUTH_MISSING两种模式不可互换——预检结果会告知您应使用哪种调用形式。
AUTH_OK- — 环境中存在令牌。本技能中的HTTP调用可直接使用。
AUTH_OK curl - — Apify CLI保存会话,且令牌无法从磁盘读取。
AUTH_OK cli: <prefix>仅包含账户元数据(用户名、套餐、代理组——无~/.apify/auth.json字段);当前CLI版本会将令牌存储在操作系统的密钥管理后端。请勿尝试从该文件提取令牌:伪造的token标头会返回Authorization: Bearer,看起来与令牌撤销的情况完全一致。请改用401命令,并在前面加上预检结果中apify call后的前缀(cli:,或当CLI不在apify中时使用PATH)。npx --yes apify-cli@latest
If the preflight prints AUTH_OK cli
AUTH_OK cli如果预检结果显示AUTH_OK cli
AUTH_OK cliEvery payload in this skill still applies — write it to a file and hand it to
instead of :
apify callcurlbash
cat > /tmp/ig-input.json <<'EOF'
{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"details","resultsLimit":1}
EOF
apify call apify/instagram-scraper --input-file /tmp/ig-input.json --output-dataset --silent- Always pass the Actor id explicitly. With no id, runs the Actor defined by a local
apify call— inside an Actor repo that silently runs the wrong thing..actor/actor.json - Prefer over inline
--input-file. Inline JSON has to survive the shell, and PowerShell and-i '{...}'mangle the quoting. A file never does.cmd.exereads stdin.--input-file - - prints the dataset to stdout — the same array
--output-datasetreturns, so every field in the Output section below is identical.run-sync-get-dataset-itemskeeps run logs off stdout so the output parses as JSON.--silent - waits for the run to finish, so the async polling pattern below is only needed on the
apify callpath. Addcurlto bound a long job.--timeout <seconds> - The run is billed to whichever account the CLI is logged in as. prints it — worth showing the user if they may have more than one.
apify info
本技能中的所有请求参数仍然适用——将其写入文件,然后通过命令执行,而非:
apify callcurlbash
cat > /tmp/ig-input.json <<'EOF'
{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"details","resultsLimit":1}
EOF
apify call apify/instagram-scraper --input-file /tmp/ig-input.json --output-dataset --silent- **请始终显式传递Actor ID。**如果不指定ID,会运行本地
apify call中定义的Actor——在Actor仓库中可能会静默执行错误的任务。.actor/actor.json - **优先使用而非内联
--input-file。**内联JSON需要在shell中正确解析,而PowerShell和-i '{...}'会破坏引号格式。使用文件则不会出现此问题。cmd.exe表示从标准输入读取内容。--input-file - - 会将数据集打印到标准输出——与
--output-dataset返回的数组完全相同,因此下文输出部分的所有字段都是一致的。run-sync-get-dataset-items参数会阻止运行日志输出到标准输出,确保输出内容可解析为JSON。--silent - 会等待任务完成,因此仅在使用
apify call方式时才需要下文的异步轮询模式。可添加curl参数限制长任务的运行时间。--timeout <seconds> - 任务费用将计入CLI登录的账户。会显示该账户信息——如果用户可能拥有多个账户,建议告知用户。
apify info
If the preflight prints AUTH_MISSING
AUTH_MISSING如果预检结果显示AUTH_MISSING
AUTH_MISSINGOpen the Apify sign-up page in the user's browser, then ask for the token. Run this
exactly as-is — it picks the right command per platform and degrades to printing the
URL when there is no browser (CI, SSH, containers):
bash
URL="https://console.apify.com/sign-up?fpr=z8j1nz"在用户浏览器中打开Apify注册页面,然后请求令牌。请严格按如下命令执行——它会根据平台选择正确的命令,当无法打开浏览器时(CI、SSH、容器环境)会回退为打印URL:
bash
URL="https://console.apify.com/sign-up?fpr=z8j1nz"xdg-open is checked before open: on some Linux distros open
is openvt, not a browser.
open优先检查xdg-open而非open:在部分Linux发行版中,open
是openvt命令,而非浏览器启动命令。
openif command -v xdg-open >/dev/null 2>&1; then xdg-open "$URL"
elif command -v open >/dev/null 2>&1; then open "$URL"
elif command -v powershell.exe >/dev/null 2>&1; then powershell.exe -NoProfile -Command "Start-Process '$URL'"
elif command -v cmd.exe >/dev/null 2>&1; then cmd.exe /c start "" "$URL"
else echo "Open this URL in a browser: $URL"
fi
Then tell the user what is happening — but **do not wait for a reply**:
> I've opened Apify's sign-up page in your browser. Instagram data comes through
> Apify, which is free to start — $5 of credits every month, no credit card, enough
> for around 2,000 results.
>
> Sign up there (or log in, if you already have an account). I'm opening the
> authorisation step now too — approve it and this machine stays connected, so you
> won't have to do any of this again.
>
> _The sign-up link is a referral link._
Immediately run the login. **Do not ask the user to confirm they've signed up first** —
this command opens Apify's authorisation page, which offers sign-up itself, and then
blocks until the user finishes. The command *is* the wait, so there is nothing to
detect and no round trip to burn:
```bash
timeout 300 npx --yes apify-cli@latest login -m console- is required. Bare
-m consolefirst prompts for a login method and will hang forever on stdin the agent cannot answer.apify logingoes straight to the browser flow.-m console - bounds it, so an abandoned sign-up doesn't hang the session. Exit code
timeout 300means it timed out — the user walked away, so ask whether they still want to continue rather than silently retrying.124 - On success it prints "Success: You are logged in to Apify as . Your token is stored in your OS keyring." From then on the preflight resolves to
<username>in every future session — no token to paste, store, or leak into a transcript.AUTH_OK cli
Then re-run the preflight. That is how you know whether it worked — not the
message, and not an assumption that the browser step went fine:
| Preflight now says | Meaning | Do |
|---|---|---|
| Logged in and persisted | Continue with the request |
| Login did not complete | Say so plainly and ask whether to retry — do not loop |
The preflight's CLI branch is just , which exits when a session
exists and non-zero when it doesn't. It is the same check that produced
a moment earlier, so re-running it is a genuine confirmation rather
than a restatement of what you already believed.
apify info0AUTH_MISSINGTwo failure modes worth naming, because both look like success from the outside:
the user closes the authorisation tab without approving (exit at the timeout),
and the user signs up but never reaches the approval step. In both the sign-up may
well have succeeded while this machine is still not connected — which is exactly
why the preflight, not the sign-up, is the thing to trust.
124Keep the sign-up tab and the login in that order. The sign-up page is where the
referral is recorded; the authorisation step afterwards is just this machine
connecting to whichever account now exists. Opening them the other way round loses
the attribution.
if command -v xdg-open >/dev/null 2>&1; then xdg-open "$URL"
elif command -v open >/dev/null 2>&1; then open "$URL"
elif command -v powershell.exe >/dev/null 2>&1; then powershell.exe -NoProfile -Command "Start-Process '$URL'"
elif command -v cmd.exe >/dev/null 2>&1; then cmd.exe /c start "" "$URL"
else echo "请在浏览器中打开此URL:$URL"
fi
然后告知用户当前操作——但**请勿等待回复**:
> 我已在您的浏览器中打开Apify的注册页面。Instagram数据通过Apify获取,该平台可免费开始使用——每月提供5美元额度,无需信用卡,足够获取约2000条结果。
>
> 请在该页面注册(或登录已有账户)。我现在将打开授权步骤页面——完成授权后,此设备将保持连接状态,您无需再重复这些操作。
>
> _注:该注册链接为推荐链接。_
立即运行登录命令。**请勿先询问用户是否已完成注册**——此命令会打开Apify的授权页面,该页面本身也提供注册选项,然后会阻塞直到用户完成操作。该命令本身就是等待过程,因此无需检测或进行往返通信:
```bash
timeout 300 npx --yes apify-cli@latest login -m console- **必须添加参数。**不带参数的
-m console会先提示选择登录方式,若代理无法回答则会在标准输入处永久挂起。apify login会直接进入浏览器授权流程。-m console - ****用于限制超时时间,避免因用户放弃注册导致会话挂起。退出码
timeout 300表示超时——用户已离开,此时应询问用户是否仍要继续,而非静默重试。124 - 成功登录后会打印*"Success: You are logged in to Apify as . Your token is stored in your OS keyring."*。此后,在所有未来会话中,预检结果都会显示为
<username>——无需粘贴、存储令牌,也不会在会话记录中泄露令牌。AUTH_OK cli
然后重新运行预检步骤。这是确认登录是否成功的唯一方式——不要依赖提示信息,也不要假设浏览器步骤已成功完成:
| 预检当前结果 | 含义 | 操作 |
|---|---|---|
| 已登录且会话已持久化 | 继续处理用户请求 |
| 登录未完成 | 明确告知用户,并询问是否重试——请勿循环执行 |
预检的CLI分支仅执行命令,当会话存在时返回退出码,不存在时返回非零值。这与之前产生的检查逻辑相同,因此重新运行预检是真实的确认,而非重复已有的假设。
apify info0AUTH_MISSING有两种值得注意的失败模式,从外部看起来像是成功:用户未批准就关闭授权标签页(超时后退出码为),以及用户完成注册但未进入授权步骤。这两种情况下,注册可能已成功,但此设备仍未连接到账户——这正是为什么要信任预检结果,而非注册流程的原因。
124**请保持先打开注册页面、再执行登录命令的顺序。**注册页面会记录推荐信息;后续的授权步骤只是将此设备连接到已创建的账户。如果顺序颠倒,会丢失推荐归因。
Never ask the user for their token
切勿向用户索要令牌
Do not request, accept, or handle an Apify API token in the conversation. The
browser login above exists precisely so the secret never reaches the agent: the CLI
receives it directly from Apify and writes it to the OS keyring, and this skill only
ever reads the result of that ('s exit code), never the value.
apify infoIf the user offers a token unprompted, decline it and point them at one of the two
safe routes below. Anything pasted into a session is a live credential sitting in a
transcript, in scrollback, and in any log that captures the conversation.
Headless environments — CI, SSH, containers, anywhere the OAuth round trip cannot
open a browser. The user sets the credential themselves, out of band, before
starting the agent:
bash
undefined**请勿在对话中请求、接受或处理Apify API令牌。**上述浏览器登录流程的存在正是为了确保密钥不会传递给代理:CLI直接从Apify接收令牌并写入操作系统密钥环,本技能仅读取该操作的结果(的退出码),绝不会读取令牌值。
apify info如果用户主动提供令牌,**请拒绝并引导他们使用以下两种安全方式之一。**粘贴到会话中的任何内容都是有效的凭证,会保存在会话记录、回滚缓存和任何捕获对话的日志中。
无头环境——CI、SSH、容器等无法打开浏览器进行OAuth往返验证的环境。用户需在启动代理之前,自行在外部设置凭证:
bash
undefinedthe user runs this in their own shell / CI secret store — not via the agent
用户需在自己的shell/CI密钥存储中运行此命令——不要通过代理执行
export APIFY_TOKEN="…"
The preflight then reports `AUTH_OK curl` and everything works, with the value never
passing through the conversation.
**Or the CLI's own prompt**, which reads the token from stdin rather than the
conversation:
```bash
npx --yes apify-cli@latest login -m manualAvoid . Passing a secret as a command-line argument exposes it in
the process list to every other process on the machine, and in shell history. It also
clears the stored session before validating, so a typo or a stale value logs the
user out of a session that was working.
login -t <token>When a token does legitimately exist in the environment, always reference it as
and let the shell expand it — as every example in this skill does.
Never substitute the literal value into a command, a log line, or a message.
$APIFY_TOKENexport APIFY_TOKEN="…"
此时预检结果会显示`AUTH_OK curl`,所有操作均可正常进行,且令牌值不会通过对话传递。
**或使用CLI自带的提示**,该提示会从标准输入读取令牌,而非通过对话:
```bash
npx --yes apify-cli@latest login -m manual避免使用。将密钥作为命令行参数传递会在进程列表中暴露给机器上的其他进程,同时也会记录在shell历史中。此外,该命令会在验证前清除已存储的会话,因此输入错误或过期值会导致用户退出已正常工作的会话。
login -t <token>当环境中确实存在令牌时,请始终以的形式引用它,让shell自动展开——如本技能中的所有示例所示。切勿将字面值替换到命令、日志行或消息中。
$APIFY_TOKENFetching data
数据抓取
One call, synchronous, returns the items directly. Good for anything that finishes
inside ~60 s. This is the form — under , put the
same payload in a file and run it through as shown above.
AUTH_OK curlAUTH_OK cli-dapify callbash
curl -s -X POST \
"https://api.apify.com/v2/acts/apify~instagram-scraper/run-sync-get-dataset-items?timeout=120" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-H "User-Agent: instagram-scraper-skill" \
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"details","resultsLimit":1}'The header is only so runs originating from this skill can be told apart
in logs. It carries no personal data and can be removed.
User-Agent单次同步调用,直接返回数据项。适用于任何可在约60秒内完成的任务。以下为模式的调用示例——若为模式,请将相同的参数写入文件,然后按上文所示通过命令执行。
AUTH_OK curlAUTH_OK cli-dapify callbash
curl -s -X POST \
"https://api.apify.com/v2/acts/apify~instagram-scraper/run-sync-get-dataset-items?timeout=120" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-H "User-Agent: instagram-scraper-skill" \
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"details","resultsLimit":1}'User-AgentChoosing resultsType
resultsType选择resultsType
resultsType| Value | Returns |
|---|---|
| Profile metadata: followers, following, bio, post count, profile picture. Cheapest — 1 result per profile. |
| A feed of posts |
| Reels only |
| Currently-live stories |
| Comments on a post URL |
| Posts mentioning an account |
Pick whenever the user only needs profile numbers. Using for that
question costs up to 100× more for no benefit.
detailsposts| 值 | 返回内容 |
|---|---|
| 主页元数据:粉丝数、关注数、简介、帖子数、头像。成本最低——每个主页返回1条结果。 |
| 帖子动态流 |
| 仅返回Reels内容 |
| 当前正在发布的快拍 |
| 帖子URL对应的评论 |
| 提及该账户的帖子 |
当用户仅需要主页数据时,请选择。使用来获取此类信息会导致成本增加多达100倍,且无任何额外收益。
detailspostsInput reference
参数参考
| Field | Type | Default | Notes |
|---|---|---|---|
| array | — | Profile, post, reel, hashtag, location or audio URLs — see URL handling below |
| string | | See table above |
| integer | | Per URL. Drives cost — always set it. |
| string | — | |
| string | — | Keyword, instead of |
| string | | |
| integer | | Max items discovered per search |
| boolean | | Stamps each item with the query that produced it |
Four more fields work but are not in the published input schema. They are
documented in the Actor's README and verified working — use them freely:
| Field | Type | Applies to | Notes |
|---|---|---|---|
| boolean | | Adds a |
| boolean | | Exclude pinned posts |
| boolean | | Newest-first ordering. Paid plans only — free plans get default order |
| boolean | | Include replies. Paid plans only. Each reply is a separate result, so totals exceed |
| 字段 | 类型 | 默认值 | 说明 |
|---|---|---|---|
| 数组 | — | 主页、帖子、Reels、话题标签、地点或音频URL——详见下文URL处理部分 |
| 字符串 | | 见上表 |
| 整数 | | 每个URL对应的结果数量。直接影响成本——请务必设置此参数。 |
| 字符串 | — | |
| 字符串 | — | 关键词,替代 |
| 字符串 | | |
| 整数 | | 每次搜索最多发现的项目数量 |
| 布尔值 | | 为每个数据项添加生成它的查询信息 |
**还有四个字段可正常使用,但未包含在已发布的参数 schema 中。**这些字段已在Actor的README中记录并验证可用——可自由使用:
| 字段 | 类型 | 适用场景 | 说明 |
|---|---|---|---|
| 布尔值 | | 添加 |
| 布尔值 | | 排除置顶帖子 |
| 布尔值 | | 按最新顺序排序评论。仅付费套餐可用——免费套餐使用默认排序 |
| 布尔值 | | 包含回复内容。仅付费套餐可用。每条回复视为单独结果,因此总数可能超过 |
URL handling
URL处理
directUrls- Profile IDs work anywhere a profile URL does — a bare numeric ID is fine
- —
instagram.com/_u/natgeo/profilecard/and_uare strippedprofilecard - — reduced to the username
instagram.com/stories/username/ - — resolved to the canonical post URL
instagram.com/share/BAC6cDeb_- - — the ID alone is valid, no slug needed
instagram.com/explore/locations/7538318/
Not supported: numeric post IDs in URL form ()
for , , or . Use the shortCode form instead. This
format does work for .
instagram.com/p/3369450800358839406/postsreelsmentionsdetailscommentsThe URL type drives the output schema. Hashtag, location, audio and explore URLs
return their own metadata even when paired with another content mode — so a location
URL with yields place details, not profile details.
resultsType: "details"directUrls- 主页ID可替代主页URL使用——纯数字ID即可
- —
instagram.com/_u/natgeo/profilecard/和_u会被自动移除profilecard - — 会简化为用户名
instagram.com/stories/username/ - — 会解析为标准帖子URL
instagram.com/share/BAC6cDeb_- - — 仅ID即可,无需slug
instagram.com/explore/locations/7538318/
不支持:在、、或模式下使用数字帖子ID格式的URL()。请改用短代码格式。该格式在模式下可以使用。
postsreelsmentionsdetailsinstagram.com/p/3369450800358839406/comments**URL类型决定输出 schema。**话题标签、地点、音频和探索URL会返回各自的元数据,即使与其他内容模式搭配使用也是如此——例如,使用的地点URL会返回地点详情,而非主页详情。
resultsType: "details"Constraints that will bite you
需要注意的限制
- One content type per run. There is no way to get posts and comments in a single call. Run twice.
- URLs beat search. and
directUrlscannot be combined; if both are present the URLs win and the search is ignored.search - Hashtags go in as plaintext — , never
travel.#travel - Multiple search terms are comma-separated in one string: .
"travel, fitness" - Free plans get about one page of comments per post (~15). Paid plans have no such cap. Do not report this as an error — say what it is.
- **每次运行仅支持一种内容类型。**无法在单次调用中同时获取帖子和评论。需运行两次调用。
- URL优先级高于搜索。和
directUrls无法同时使用;如果两者都存在,URL会优先生效,搜索参数会被忽略。search - 话题标签需以纯文本形式传入——,而非
travel。#travel - 多个搜索词用逗号分隔为单个字符串:。
"travel, fitness" - **免费套餐每个帖子最多获取约一页评论(约15条)。**付费套餐无此限制。请勿将此情况报告为错误——如实告知用户即可。
Common tasks
常见任务
Every capability the Actor exposes, with the payload for each. Swap the
into the curl above.
-d '...'Profile stats for several accounts — cheapest possible call:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/","https://www.instagram.com/natgeo/"],"resultsType":"details","resultsLimit":1}'A profile's recent posts, last 30 days:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"posts","resultsLimit":30,"onlyPostsNewerThan":"1 month"}'A profile's reels:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"reels","resultsLimit":20}'A profile's current stories — only returns anything while stories are live, and
often needs a paid plan:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"stories","resultsLimit":20}'Posts that mention an account — brand monitoring:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"mentions","resultsLimit":50}'Comments on one post:
bash
-d '{"directUrls":["https://www.instagram.com/p/SHORTCODE/"],"resultsType":"comments","resultsLimit":50}'Posts under a hashtag:
bash
-d '{"search":"wildlifephotography","searchType":"hashtag","searchLimit":50,"resultsType":"posts","resultsLimit":50}'Find accounts by keyword — returns account records:
userbash
-d '{"search":"climate photographer","searchType":"user","searchLimit":20,"resultsType":"details"}'Search profiles by name — matches on profile pages:
profilebash
-d '{"search":"national geographic","searchType":"profile","searchLimit":20,"resultsType":"details"}'Posts from a place:
bash
-d '{"search":"Yosemite National Park","searchType":"place","searchLimit":20,"resultsType":"posts","resultsLimit":50}'Posts from a specific location or hashtag URL — pass the URL directly instead of
searching:
bash
-d '{"directUrls":["https://www.instagram.com/explore/tags/wildlife/"],"resultsType":"posts","resultsLimit":50}'Tracking which query produced which post — when scraping several hashtags or
profiles in one run, stamps each item with its source so the results
can be grouped afterwards:
addParentDatabash
-d '{"search":"wildlife","searchType":"hashtag","searchLimit":30,"resultsType":"posts","resultsLimit":30,"addParentData":true}'Deep profile statistics — account type, post and reel counts, category, city,
public contact fields. Also works on private profiles:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"details","resultsLimit":1,"addProfileStatistics":true}'Recent posts, excluding pinned ones:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"posts","resultsLimit":30,"skipPinnedPosts":true}'Only the pinned posts — invert the same pair:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"posts","skipPinnedPosts":false,"onlyPostsNewerThan":"0 minutes"}'Newest comments first, including replies — both paid-plan only:
bash
-d '{"directUrls":["https://www.instagram.com/p/SHORTCODE/"],"resultsType":"comments","resultsLimit":50,"isNewestComments":true,"includeNestedComments":true}'以下为Actor支持的所有功能及对应的请求参数。将部分替换到上述curl命令中即可。
-d '...'多个账户的主页统计数据——成本最低的调用方式:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/","https://www.instagram.com/natgeo/"],"resultsType":"details","resultsLimit":1}'主页的近期帖子(过去30天):
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"posts","resultsLimit":30,"onlyPostsNewerThan":"1 month"}'主页的Reels内容:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"reels","resultsLimit":20}'主页的当前快拍——仅当快拍处于发布状态时才会返回数据,且通常需要付费套餐:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"stories","resultsLimit":20}'提及该账户的帖子——品牌监控:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"mentions","resultsLimit":50}'单条帖子的评论:
bash
-d '{"directUrls":["https://www.instagram.com/p/SHORTCODE/"],"resultsType":"comments","resultsLimit":50}'特定话题标签下的帖子:
bash
-d '{"search":"wildlifephotography","searchType":"hashtag","searchLimit":50,"resultsType":"posts","resultsLimit":50}'按关键词查找账户——类型返回账户记录:
userbash
-d '{"search":"climate photographer","searchType":"user","searchLimit":20,"resultsType":"details"}'按名称搜索主页——类型匹配主页页面:
profilebash
-d '{"search":"national geographic","searchType":"profile","searchLimit":20,"resultsType":"details"}'特定地点的帖子:
bash
-d '{"search":"Yosemite National Park","searchType":"place","searchLimit":20,"resultsType":"posts","resultsLimit":50}'特定地点或话题标签URL下的帖子——直接传入URL而非使用搜索:
bash
-d '{"directUrls":["https://www.instagram.com/explore/tags/wildlife/"],"resultsType":"posts","resultsLimit":50}'跟踪每个查询对应的帖子——当在单次运行中抓取多个话题标签或主页时,会为每个数据项添加来源信息,以便后续对结果进行分组:
addParentDatabash
-d '{"search":"wildlife","searchType":"hashtag","searchLimit":30,"resultsType":"posts","resultsLimit":30,"addParentData":true}'详细主页统计数据——账户类型、帖子和Reels数量、分类、城市、公开联系字段。对私人主页也有效:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"details","resultsLimit":1,"addProfileStatistics":true}'近期帖子(排除置顶帖子):
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"posts","resultsLimit":30,"skipPinnedPosts":true}'仅获取置顶帖子——反转上述参数组合:
bash
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"posts","skipPinnedPosts":false,"onlyPostsNewerThan":"0 minutes"}'最新评论优先,包含回复内容——均仅付费套餐可用:
bash
-d '{"directUrls":["https://www.instagram.com/p/SHORTCODE/"],"resultsType":"comments","resultsLimit":50,"isNewestComments":true,"includeNestedComments":true}'Output
输出
Each content type produces a different schema, and they cannot be combined in one
run. Samples below are abridged to the useful fields; the URL type you pass can
override the shape (a location URL returns place data even under another mode).
Post / carousel () — 23 fields:
resultsType: "posts"json
{
"inputUrl": "https://www.instagram.com/p/DZxvMgyH8yR/",
"id": "3923124318436838545",
"type": "Image",
"shortCode": "DZxvMgyH8yR",
"caption": "In the Democratic Republic of the Congo…",
"hashtags": [], "mentions": ["carstenpeter"],
"url": "https://www.instagram.com/p/DZxvMgyH8yR/",
"likesCount": 34512, "commentsCount": 210,
"timestamp": "2026-06-22T17:24:20.000Z",
"displayUrl": "https://…", "images": [], "childPosts": [],
"dimensionsWidth": 1080, "dimensionsHeight": 1350, "alt": "…",
"ownerUsername": "natgeo", "ownerFullName": "National Geographic", "ownerId": "787132",
"firstComment": "…", "latestComments": [], "isCommentsDisabled": false
}Reel () — 31 fields. Posts plus video:
resultsType: "reels"json
{
"…all post fields…": "…",
"videoUrl": "https://…", "videoDuration": 28.28,
"videoViewCount": 120433, "videoPlayCount": 98221,
"audioUrl": "https://…", "musicInfo": {}, "productType": "clips",
"isPinned": false
}Comment ():
resultsType: "comments"json
{
"postUrl": "https://www.instagram.com/p/DZ5T2XPllXv/",
"commentUrl": "https://www.instagram.com/p/DZ5T2XPllXv/c/18093536360613690",
"id": "18093536360613690",
"text": "We love you NASA 💙🌎🌊",
"ownerUsername": "mavideniz5521__", "ownerProfilePicUrl": "https://…",
"timestamp": "2026-06-22T17:24:20.000Z",
"likesCount": 4, "repliesCount": null, "replies": null,
"owner": { "username": "…" }
}repliesCountrepliesincludeNestedComments: trueProfile details () — 22 fields. Verified live:
resultsType: "details"json
{
"inputUrl": "https://www.instagram.com/nasa/",
"id": "528817151", "username": "nasa", "url": "https://www.instagram.com/nasa/",
"fullName": "NASA", "biography": "Making the seemingly impossible, possible. ✨",
"externalUrl": "https://www.nasa.gov/", "externalUrls": [],
"followersCount": 104423132, "followsCount": 92, "postsCount": 4888,
"verified": true, "private": false,
"isBusinessAccount": true, "businessCategoryName": "Government Agencies",
"joinedRecently": false, "fbid": "17841401474538262",
"profilePicUrl": "https://…", "profilePicUrlHD": "https://…",
"highlightReelCount": 5, "igtvVideoCount": 171,
"latestPosts": [], "relatedProfiles": []
}detailslatestPostsrelatedProfilesdetailspostsWith a object is appended (~60 fields):
(1 Personal, 2 Business, 3 Creator), , ,
, , , , , .
addProfileStatistics: truestatisticsaccount_typemedia_counttotal_clips_countcategorycity_nameaddress_streetzipbio_linksmutual_followers_countMentions () — 21 fields, post-shaped, plus ,
, , .
resultsType: "mentions"taggedUsersmusiccarouselImagescarouselImageCountPlace details (location URL) — 16 fields:
json
{
"inputUrl": "https://www.instagram.com/explore/locations/7538318/",
"name": "Copenhagen, Denmark", "location_id": "7538318", "slug": "copenhagen",
"lat": 55.6761, "lng": 12.5683,
"location_address": "…", "location_city": "…", "location_zip": "…",
"phone": "…", "category": "…", "price_range": "…",
"media_count": 1284322, "ig_business": "…", "posts": [], "hours": {}
}Hashtag details (hashtag URL) — 15 fields, including SEO-style extras:
, , , , , , , ,
, , , , , .
namepostsCounturlidpostspostsPerDaydifficultyrelatedfrequentaveragerarerelatedFrequentrelatedAveragerelatedRareSearch results carry and so you can tell which query
produced each row:
searchTermsearchSource- Hashtag search — 6 fields: ,
searchTerm,searchSource,name,postsCount,urlid - Place search — 17 fields: place-details shape plus /
searchTermsearchSource - Profile search — 13 fields: post-shaped, with ,
ownerUsername,ownerFullNametaggedUsers
**每种内容类型对应不同的schema,且无法在单次运行中混合使用。**以下示例仅保留有用字段;传入的URL类型可能会覆盖输出格式(例如,地点URL即使搭配其他模式也会返回地点数据)。
帖子/轮播帖()——共23个字段:
resultsType: "posts"json
{
"inputUrl": "https://www.instagram.com/p/DZxvMgyH8yR/",
"id": "3923124318436838545",
"type": "Image",
"shortCode": "DZxvMgyH8yR",
"caption": "In the Democratic Republic of the Congo…",
"hashtags": [], "mentions": ["carstenpeter"],
"url": "https://www.instagram.com/p/DZxvMgyH8yR/",
"likesCount": 34512, "commentsCount": 210,
"timestamp": "2026-06-22T17:24:20.000Z",
"displayUrl": "https://…", "images": [], "childPosts": [],
"dimensionsWidth": 1080, "dimensionsHeight": 1350, "alt": "…",
"ownerUsername": "natgeo", "ownerFullName": "National Geographic", "ownerId": "787132",
"firstComment": "…", "latestComments": [], "isCommentsDisabled": false
}Reels()——共31个字段。包含帖子的所有字段,外加视频相关字段:
resultsType: "reels"json
{
"…all post fields…": "…",
"videoUrl": "https://…", "videoDuration": 28.28,
"videoViewCount": 120433, "videoPlayCount": 98221,
"audioUrl": "https://…", "musicInfo": {}, "productType": "clips",
"isPinned": false
}评论():
resultsType: "comments"json
{
"postUrl": "https://www.instagram.com/p/DZ5T2XPllXv/",
"commentUrl": "https://www.instagram.com/p/DZ5T2XPllXv/c/18093536360613690",
"id": "18093536360613690",
"text": "We love you NASA 💙🌎🌊",
"ownerUsername": "mavideniz5521__", "ownerProfilePicUrl": "https://…",
"timestamp": "2026-06-22T17:24:20.000Z",
"likesCount": 4, "repliesCount": null, "replies": null,
"owner": { "username": "…" }
}仅当设置(付费套餐)时,和字段才会填充数据。
includeNestedComments: truerepliesCountreplies主页详情()——共22个字段。已验证可用:
resultsType: "details"json
{
"inputUrl": "https://www.instagram.com/nasa/",
"id": "528817151", "username": "nasa", "url": "https://www.instagram.com/nasa/",
"fullName": "NASA", "biography": "Making the seemingly impossible, possible. ✨",
"externalUrl": "https://www.nasa.gov/", "externalUrls": [],
"followersCount": 104423132, "followsCount": 92, "postsCount": 4888,
"verified": true, "private": false,
"isBusinessAccount": true, "businessCategoryName": "Government Agencies",
"joinedRecently": false, "fbid": "17841401474538262",
"profilePicUrl": "https://…", "profilePicUrlHD": "https://…",
"highlightReelCount": 5, "igtvVideoCount": 171,
"latestPosts": [], "relatedProfiles": []
}**模式已包含(最多12条)和(最多48个),且无需额外成本。**如果用户需要主页数据和近期帖子预览,一次调用即可满足需求——通常无需进行第二次调用。
detailslatestPostsrelatedProfilesdetailsposts设置后,会追加一个对象(约60个字段):(1=个人、2=企业、3=创作者)、、、、、、、、。
addProfileStatistics: truestatisticsaccount_typemedia_counttotal_clips_countcategorycity_nameaddress_streetzipbio_linksmutual_followers_count提及内容()——共21个字段,格式与帖子类似,外加、、、字段。
resultsType: "mentions"taggedUsersmusiccarouselImagescarouselImageCount地点详情(地点URL)——共16个字段:
json
{
"inputUrl": "https://www.instagram.com/explore/locations/7538318/",
"name": "Copenhagen, Denmark", "location_id": "7538318", "slug": "copenhagen",
"lat": 55.6761, "lng": 12.5683,
"location_address": "…", "location_city": "…", "location_zip": "…",
"phone": "…", "category": "…", "price_range": "…",
"media_count": 1284322, "ig_business": "…", "posts": [], "hours": {}
}话题标签详情(话题标签URL)——共15个字段,包括类SEO的额外字段:、、、、、、、、、、、、、。
namepostsCounturlidpostspostsPerDaydifficultyrelatedfrequentaveragerarerelatedFrequentrelatedAveragerelatedRare搜索结果包含和字段,可用于区分每个结果对应的查询:
searchTermsearchSource- 话题标签搜索——6个字段:、
searchTerm、searchSource、name、postsCount、urlid - 地点搜索——17个字段:地点详情格式外加/
searchTerm字段searchSource - 主页搜索——13个字段:帖子格式,包含、
ownerUsername、ownerFullName字段taggedUsers
Images are URLs, not files
图片为URL,而非文件
Every image and video field is a link to Instagram's CDN. Nothing is downloaded, and
those links are signed and expire after a few hours. If the user needs the media
itself, fetch it promptly on their own bandwidth.
所有图片和视频字段均为指向Instagram CDN的链接。不会下载任何文件,且这些链接已签名,几小时后会过期。如果用户需要媒体文件本身,请尽快使用自己的带宽下载。
Runs longer than a minute
运行时间超过一分钟的任务
For large jobs, start async and poll rather than holding a sync connection:
bash
RUN=$(curl -s -X POST "https://api.apify.com/v2/acts/apify~instagram-scraper/runs" \
-H "Authorization: Bearer $APIFY_TOKEN" -H "Content-Type: application/json" \
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"posts","resultsLimit":1000}' \
| python3 -c "import sys,json; print(json.load(sys.stdin)['data']['id'])")
curl -s "https://api.apify.com/v2/actor-runs/$RUN?waitForFinish=60" \
-H "Authorization: Bearer $APIFY_TOKEN" \
| python3 -c "import sys,json; print(json.load(sys.stdin)['data']['status'])"
curl -s "https://api.apify.com/v2/actor-runs/$RUN/dataset/items?format=json" \
-H "Authorization: Bearer $APIFY_TOKEN"Poll until is , then fetch items. or means stop
and report — do not silently retry a job the user is paying for.
statusSUCCEEDEDFAILEDABORTED对于大型任务,请启动异步调用并进行轮询,而非保持同步连接:
bash
RUN=$(curl -s -X POST "https://api.apify.com/v2/acts/apify~instagram-scraper/runs" \
-H "Authorization: Bearer $APIFY_TOKEN" -H "Content-Type: application/json" \
-d '{"directUrls":["https://www.instagram.com/nasa/"],"resultsType":"posts","resultsLimit":1000}' \
| python3 -c "import sys,json; print(json.load(sys.stdin)['data']['id'])")
curl -s "https://api.apify.com/v2/actor-runs/$RUN?waitForFinish=60" \
-H "Authorization: Bearer $APIFY_TOKEN" \
| python3 -c "import sys,json; print(json.load(sys.stdin)['data']['status'])"
curl -s "https://api.apify.com/v2/actor-runs/$RUN/dataset/items?format=json" \
-H "Authorization: Bearer $APIFY_TOKEN"轮询直到变为,然后获取数据项。如果为或,请停止并报告——请勿静默重试用户需要付费的任务。
statusSUCCEEDEDstatusFAILEDABORTEDErrors
错误处理
| Status | Meaning | What to do |
|---|---|---|
| Token missing, wrong, or revoked | Re-run the Setup preflight above. Do not retry the call. A |
| Monthly credits exhausted | Tell the user; they can wait for the monthly reset or upgrade at https://apify.com/pricing?fpr=z8j1nz (referral link). Do not retry. |
| Bad Actor ID or run ID | Check the URL uses |
| Sync call exceeded the limit | Switch to the async pattern above. |
| Empty array | Private, deleted, or genuinely empty | Report honestly. Do not assume it is a billing problem. |
| The run failed, or the CLI session is gone | The CLI prints the reason and a run URL — read it rather than retrying. If |
An empty result is a real answer. Private accounts, deleted posts and quiet hashtags
all legitimately return nothing.
| 状态码 | 含义 | 操作 |
|---|---|---|
| 令牌缺失、错误或已撤销 | 重新运行上述设置中的预检步骤。请勿重试调用。如果在curl模式下收到 |
| 每月额度已耗尽 | 告知用户;他们可以等待每月额度重置,或访问https://apify.com/pricing?fpr=z8j1nz(推荐链接)升级套餐。请勿重试。 |
| Actor ID或运行ID错误 | 检查URL是否使用 |
| 同步调用超出时间限制 | 切换为上述异步模式。 |
| 空数组 | 内容为私人、已删除或确实为空 | 如实报告。请勿假设这是计费问题。 |
| 任务失败或CLI会话已失效 | CLI会打印原因和运行URL——请查看该信息而非重试。如果 |
空结果是真实的返回值。私人账户、已删除帖子和无内容的话题标签都会合法地返回空结果。
Data quirks to report accurately, not treat as bugs
需要准确报告的数据异常,而非视为错误
- means the creator hid the like count. Instagram does not expose it. Say "hidden by the creator" — never report it as zero or as an error.
likesCount: -1 - Private profiles generally return nothing. One exception: if a private account is tagged as a collaborator on a post and any co-author is public, Instagram treats that post as public and it will appear in results.
- Metrics can differ from what the app shows. Instagram serves slightly different counts to logged-out visitors, and large counts move constantly. Small discrepancies are expected.
- Result counts are not guaranteed. There is no fixed cap; you get what Instagram exposes publicly. To sanity-check what should be available, open the URL in an incognito window.
- ****表示创作者隐藏了点赞数。Instagram不会公开该数据。请告知用户“点赞数已被创作者隐藏”——切勿报告为0或错误。
likesCount: -1 - 私人主页通常不会返回任何数据。例外情况:如果私人账户作为合作者被标记在帖子中,且任何共同作者为公开账户,Instagram会将该帖子视为公开内容,会出现在结果中。
- **指标可能与应用显示的不同。**Instagram向未登录访客提供的计数略有不同,且大数据量的计数会持续变化。存在微小差异是正常的。
- **结果数量不保证。**没有固定上限;您只能获取Instagram公开的内容。要验证应获取的内容,可在隐身窗口中打开对应的URL。
Cost discipline
成本控制
The user is paying per result. Treat that as real money.
- Always set . The default is 100 per URL; a five-URL call with the default costs 500 results when the user probably wanted 25.
resultsLimit - Use for any question about followers, bio or post counts. It returns one result per profile.
resultsType: "details" - Start small. For anything open-ended, run a bounded first pass, show the user what came back, and confirm before scaling up.
- Say what a large run will cost before starting it. At $2.30/1,000, a 5,000-result job is about $11.50.
- Never loop the same call after a or
401— each attempt can be billable and none of them will succeed.402
用户按结果数量付费。请将其视为真实货币。
- **请始终设置。**默认值为每个URL返回100条结果;如果用户需要25条结果,使用默认值的5个URL调用会消耗500条结果。
resultsLimit - 对于粉丝数、简介或帖子数相关的查询,请使用。每个主页仅返回1条结果。
resultsType: "details" - **从小规模开始。**对于任何开放式查询,请先运行有限制的测试调用,向用户展示结果,确认后再扩大规模。
- **在启动大型任务前告知用户成本。**按每1000条结果2.30美元计算,5000条结果的任务成本约为11.50美元。
- 在收到或
401错误后,切勿重复调用相同任务——每次尝试都可能产生费用,且不会成功。402