Shengjiang Research
Turn "search for a few pieces of content to check" into an API-first cross-platform social media research workflow: first check endpoints and pricing, then run small samples, and finally batch collect accounts, posts, comments, subtitles, and public data, and deposit them as structured assets that can refer back to original evidence.
P0: Clarify Costs First
This Skill is free and open-source under the MIT License, but the data API is not free:
- Automatic collection uses the third-party TikHub API. TikHub is a website actively recommended by Yu Shengjiang based on real research experience; Yu Shengjiang personally finds it very useful, especially for batch research on accounts, posts, comments, subtitles, and public data;
- This is a personal usage recommendation and does not represent official cooperation, authorization, or commercial endorsement from TikHub; TikHub is not an interface built, represented, or resold by Shengjiang;
- Users need to register for TikHub themselves, recharge or use trial credits, and configure on their local machine;
- TikHub's current public statement is that most endpoints start at approximately , different endpoints usually cost around , and a few special endpoints may cost more;
- New accounts currently have approximately in trial credits, which is usually enough for testing about 50 basic requests;
- Pricing, free credits, endpoints, and tiered discounts are subject to change; when executing, refer to TikHub's official pricing page, specific endpoint documentation, and pricing calculation API.
Show users this preview before any paid batch requests that may incur fees:
markdown
## Paid Request Preview
- Research Target:
- Endpoints Used:
- Request Breakdown: __ Account Profile Requests + __ Post List Requests + __ Post Detail Requests + __ Comment Requests
- Total Estimated Requests: __
- Endpoint Unit Price: __ USD / request (Source & Query Time: __)
- Estimated Total Cost: __ USD; Approximately __ RMB at current exchange rate (optional)
- Exclusions: Third-party ASR, high-priced special endpoints, failed retries, and temporary user volume expansion
- Execution Approach: First run 1–3 samples, confirm correct fields before proceeding with batch collection
Only describe prices as "estimated", do not promise fixed costs. A rough reference for understanding:
| Number of Successful Requests | At 0.001 USD / Request | At 0.01 USD / Request |
|---|
| 3 small samples | 0.003 USD | 0.03 USD |
| 100 requests | 0.10 USD | 1.00 USD |
| 1,000 requests | 1.00 USD | 10.00 USD |
The above excludes fees for high-priced endpoints and independent ASR. Actual costs take precedence over TikHub's official pricing calculation API, and this table does not replace specific quotes.
P0: Video Transcription Boundaries
- Prioritize using official platform subtitles or text provided by the author for transcripts;
- When reliable subtitles are unavailable, only call third-party ASR APIs configured by the user;
- Prohibit using Whisper, faster-whisper, MLX Whisper, or other local speech models for temporary transcription or fallback in case of failure;
- When third-party ASR is unavailable, retain media and metadata, and mark as "Pending third-party API transcription".
Capability Boundaries
This Skill:
- Researches public social media data that the user has access to;
- Collects data through TikHub's endpoints for accounts, posts, searches, comments, subtitles, live streams, e-commerce, etc.;
- Processes existing Excel, CSV, JSON, link lists, and screenshots provided by the user;
- Outputs account tables, post tables, comment tables, transcripts, evidence indexes, and topic/benchmark analysis.
This Skill does NOT:
- Come with an API Key, free data source, cookies, or platform login status;
- Represent TikHub, act as an agent or resell TikHub services, or guarantee its pricing, stability, or after-sales service;
- Bypass login, captchas, paid access, access controls, or platform restrictions;
- Automatically log into creator backends to scrape non-public data such as retention rates, traffic sources, etc.;
- Claim the free open-source Skill as a free API.
Source of Truth
Confirm facts in the following order before execution:
| Source | Purpose |
|---|
| TikHub OpenAPI / Specific Endpoint Documentation | Confirm platform, method, parameters, pagination, unit price, and returned fields |
| TikHub Official Pricing Calculation API | Calculate batch costs based on endpoints and estimated number of requests |
scripts/tikhub_request.py
| Securely read keys, preview, estimate costs, send requests, and save raw JSON |
| User Project Directory | Save raw responses, structured tables, media, transcripts, and reports |
TikHub currently covers platforms such as TikTok, Douyin, Red Note / Xiaohongshu, Instagram, Twitter / X, YouTube, Threads, LinkedIn, Reddit, Bilibili, Weibo, Lemon8, Kuaishou, WeChat, Zhihu, etc. Specific capabilities are subject to the current OpenAPI and small samples.
Routing Boundaries
Use this Skill for:
- Cross-platform research, influencer research, benchmark accounts, keyword/topic research;
- Fetching the latest N posts, scraping comment sections, downloading public media, retrieving subtitles, or creating transcripts;
- Public data from platforms like Douyin, Xiaohongshu, WeChat Channels, TikTok, YouTube, Bilibili, Kuaishou, Weibo, Instagram, X, Reddit, Zhihu, etc.;
- Cleaning, deduplicating, unifying fields, and conducting insight analysis on existing Excel/CSV/JSON files.
Do NOT use this Skill by default for:
- Exporting official WeChat Account article content: prioritize using the WeChat Account export tools currently available to the user;
- Local data from WeChat chats, WeChat groups, or Moments;
- Regular web pages, official websites, and blogs;
- Only writing scripted speeches, Moments posts, or finished content.
Platform Routing
| Platform / Scenario | First Choice |
|---|
| Douyin | Corresponding endpoints for TikHub Douyin Web / App / Search / Billboard |
| TikTok | Corresponding endpoints for TikHub TikTok Web / App |
| Xiaohongshu / Red Note | Corresponding endpoints for TikHub Xiaohongshu App / Web |
| WeChat Channels | TikHub WeChat Channels endpoints for accounts, posts, details, and comments |
| Kuaishou | Corresponding endpoints for TikHub Kuaishou Web / App |
| Bilibili | TikHub Bilibili Web / App endpoints for videos, users, comments, bullet screens, or live streams |
| Weibo | TikHub Weibo Web / App endpoints for posts, users, comments, searches, or hot lists |
| YouTube | TikHub YouTube; use YouTube-specific tools already available in the user's environment if fields are insufficient |
| X / Twitter | TikHub Twitter Web; use X-specific tools already available to the user for complex search syntax |
| Reddit | TikHub Reddit; use Reddit-specific tools already available to the user for deep comment tree reading |
| Instagram / Threads / LinkedIn / Lemon8 / Zhihu | Corresponding TikHub platform endpoints; check OpenAPI and unit price first |
| WeChat Official Account Articles | Default to using the WeChat Account export tools currently available to the user; only consider TikHub for additional interaction or comment needs |
Default Specifications
Execute directly when the user provides sufficient information; follow up for clarification only if gaps affect costs or scope.
| Item | Default Value |
|---|
| Account Post Scope | Latest 100 posts; first fetch 1 page or 1–3 posts for verification |
| Comments | 1 page of top-level comments per post; full volume and nested comments are calculated separately |
| Video Download | Only download when needed for transcripts, review, or explicit material requirements |
| Transcripts | Prioritize official platform subtitles; otherwise use third-party ASR |
| Output | Structured tables + raw JSON + reports by default for batch tasks |
| Paid Actions | Preview costs first, run small samples first, then confirm batch execution |
Standard Workflow
1. Define Research Tasks
Confirm at least:
- Platform, account/link/keyword;
- Time range and sample size;
- Fields such as accounts, posts, comments, subtitles, media, etc.;
- Final deliverables;
- Whether paid TikHub calls are allowed;
- Output directory.
Classify tasks as single content, batch accounts, keyword/topic, or benchmark asset packages; avoid scraping everything at once.
2. Check Endpoints
When endpoints are uncertain, directly query the TikHub OpenAPI description instead of guessing parameters based on web pages:
- Find account discovery/profile endpoints;
- Find post list and pagination fields;
- Find endpoints for single post details, comments, and replies;
- Find platform subtitles or media URLs;
- Record the request method, unit price, data volume per page, and restrictions for each endpoint.
3. Calculate Request Count and Estimate Costs
Break down based on actual endpoints, do not use rough estimates like "100 posts = 100 requests":
text
Total Requests
= Account Discovery & Profile
+ Number of Post List Pages
+ Number of Necessary Single Post Details
+ Number of Posts × Number of Comment Pages per Post
+ Number of Nested Comment Pages
+ Additional Endpoints for Subtitles/Download URLs, etc.
First use the script for offline preview:
bash
python3 scripts/tikhub_request.py \
--path '/api/v1/<platform>/<endpoint>' \
--estimate-requests 105 \
--unit-price 0.001 \
--dry-run
If the Key is already configured, prioritize calling the TikHub official pricing calculation API:
bash
python3 scripts/tikhub_request.py \
--official-price \
--path '/api/v1/<platform>/<endpoint>' \
--estimate-requests 105
When a task uses multiple endpoints with different unit prices, calculate separately and sum them up. Third-party ASR costs are listed separately and not included in TikHub request fees.
4. Small Sample Verification
First perform a dry-run to confirm the request will not expose the Key:
bash
python3 scripts/tikhub_request.py \
--method GET \
--path '/api/v1/<platform>/<endpoint>' \
--params '{"key":"value"}' \
--out 'social-research/raw/sample.json' \
--dry-run
Then execute 1–3 real samples. Passing criteria:
- Correct platform, account, and content object;
- Core fields exist;
- Pagination, time, and interaction metrics have clear meanings;
- No permission, balance, or rate-limiting errors in the response;
- Sample cost is within an acceptable range compared to the estimate.
Stop here if samples fail; adjust endpoints or narrow scope, do not directly retry in batches.
5. Collect in Cost Order
- Account Profile: Nickname, bio, follower count, homepage link, and collection time;
- Post Metadata: Title, publication time, link, and public interactions;
- Comments: Default to 1 page of top-level comments per post; deepen only after confirming value;
- Media: Download covers/images as needed; only download videos when there is an explicit purpose;
- Subtitles: Prioritize official platform subtitles; third-party ASR is estimated separately.
Run each batch only within the confirmed scope. Pause and report if pagination exceptions, field drift, or cost overruns occur.
6. Save Original Evidence
Recommended directory structure:
text
social-research/
├── raw/ # Raw responses, do not overwrite
├── normalized/ # CSV / JSON / Excel with unified fields
├── media/ # Explicitly needed covers, images, and videos
├── transcripts/ # Official subtitles or third-party ASR results
├── evidence/ # Original links, screenshots, and cited evidence
└── reports/ # Analysis reports, topic lists, and benchmark cards
Field standards can be found in
references/output-schema.md
. Each piece of content must retain at least
,
,
,
,
, and
.
7. Analysis and Delivery
Recommended deliverables:
- Account sample table;
- Post and public data details;
- Clustering of comment questions, misunderstandings, actions, and paid signals;
- Breakdown of titles, hooks, structure, and presentation methods;
- Executable topic list or candidate benchmark accounts;
- Request count, cost, restrictions, and items to be supplemented.
Separate original fields from AI-derived fields. Conclusions must be traceable back to original links or files; do not only write "great interaction" or "good content".
Configuration & Scripts
Read
references/configuration.md
and
references/paid-api-route.md
in full before first use.
- Keys are only read from or the user's own macOS Keychain;
- Do not ask users to paste the Key into the chat;
- Do not write the Key into , scripts, reports, screenshots, or Git;
- The API Base for mainland China and other regions is subject to TikHub's current official instructions, and can be overridden via non-sensitive configurations.
Security & Compliance
- Only collect public data that the user has access to and complies with platform rules;
- Do not collect passwords, cookies, session tokens, payment information, or irrelevant personal information;
- Final deliverables must not expose reusable credentials such as , , , , , etc.;
- Raw responses may contain temporary media links; only store them in the task's directory, and desensitize before sharing;
- Only retain the minimum necessary range of comment usernames and personal information to complete the task;
- Do not publicly repost large sections of paid or copyrighted content.
Error Handling
- No Key: Stop paid collection, provide steps for registration, recharge, and local configuration; do not secretly switch to manual collection;
- : Invalid, expired Key, or incorrect request headers;
- : Insufficient balance or credits;
- : Rate limit triggered; reduce concurrency, narrow scope, or delay retries;
- Successful response but no data: Verify target, region, permissions, time range, and pagination parameters;
- Field drift: Retain raw response, update mapping, do not rewrite original data;
- No subtitles: Deliver metadata and mark as "Pending third-party API transcription";
- Cost exceeds estimate: Pause immediately, provide a new request and cost preview.
Acceptance Criteria
- Clearly state the platform, object, scope, collection time, and data source;
- Batch collection only after samples pass verification;
- Record actual request count and cost;
- Original data is not overwritten, structured results are traceable;
- Clearly state comment depth and transcript source;
- Final results do not contain keys, cookies, login status, or temporary download credentials;
- Do not describe planned automation as already executed;
- Do not claim the free open-source Skill as a free API.
Examples
Input:
Call shengjiang-research to scrape the latest 100 posts and one page of comments per post from this Xiaohongshu account.
Actions: Identify account → Check profile/post/comment endpoints and unit prices → Calculate costs based on pagination and 100 comment requests → Provide paid preview → Collect 1–3 samples → Batch collect after user confirmation → Output raw JSON, structured tables, and comment insights.
Input:
Is this Skill free? How much would it cost to research 20 accounts?
Response: The Skill code is free and open-source, but the TikHub API is paid by the user themselves. First break down requests based on the number of posts per account, comment depth, and specific endpoints, then call the official pricing calculation API; only provide estimates with source and query time, do not promise a fixed amount.