The official Node.js SDK for Scrapeless AI - End-to-End Data Infrastructure for AI Developers & Enterprises.
New to Scrapeless? Sign up and get $5 in free credits.
- π Features
- π¦ Installation
- π Quick Start
- π Usage Examples
- π§ API Reference
- π Examples
- π§ͺ Testing
- π οΈ Contributing & Development Guide
- π License
- π Support
- π’ About Scrapeless
- Browser: Advanced browser session management supporting Playwright and Puppeteer frameworks, with configurable anti-detection capabilities (e.g., fingerprint spoofing, CAPTCHA solving) and extensible automation workflows.
- Web Unlocker: web interaction and data extraction with full browser capabilities. Execute JavaScript rendering, simulate user interactions (clicks, scrolls), bypass anti-scraping measures, and export structured data in formats.
- Crawl: Extract data from single pages or traverse entire domains, exporting in formats including Markdown, JSON, HTML, screenshots, and links.
- Scraping API: Direct data extraction APIs for websites (e.g., e-commerce, travel platforms). Retrieve structured product information, pricing, and reviews with pre-built connectors.
- Google Search API: Google SERP data extraction API. Fetch organic results, news, images, and more with customizable parameters and real-time updates.
- Proxies: Geo-targeted proxy network with 195+ countries. Optimize requests for better success rates and regional data access.
- AI Scraper: Extract AI chat answers, citations, and brand mentions across supported models.
- TypeScript Support: Full TypeScript definitions for better development experience
Install the SDK using npm:
npm install @scrapeless-ai/sdkOr using yarn:
yarn add @scrapeless-ai/sdkOr using pnpm:
pnpm add @scrapeless-ai/sdkLog in to the Scrapeless Dashboard and get the API Key
import { Scrapeless } from '@scrapeless-ai/sdk';
// Initialize the client
const client = new Scrapeless({
apiKey: 'your-api-key' // Get your API key from https://scrapeless.com
});You can also configure the SDK using environment variables:
# Required
SCRAPELESS_API_KEY=your-api-key
# Optional - Custom API endpoints
SCRAPELESS_BASE_API_URL=https://api.scrapeless.com
SCRAPELESS_BROWSER_API_URL=https://browser.scrapeless.com
SCRAPELESS_CRAWL_API_URL=https://api.scrapeless.comAdvanced browser session management supporting Playwright and Puppeteer frameworks, with configurable anti-detection capabilities (e.g., fingerprint spoofing, CAPTCHA solving) and extensible automation workflows:
import { Scrapeless } from '@scrapeless-ai/sdk';
import puppeteer from 'puppeteer-core';
const client = new Scrapeless();
// Create a browser session
const { browserWSEndpoint } = await client.browser.create({
sessionName: 'my-session',
sessionTTL: 180,
proxyCountry: 'US'
});
// Connect with Puppeteer
const browser = await puppeteer.connect({
browserWSEndpoint: browserWSEndpoint
});
const page = await browser.newPage();
await page.goto('https://example.com');
console.log(await page.title());
await browser.close();Manage browser profiles for persistent sessions.
const createResponse = await client.profiles.create('My Profile');
console.log('Profile created:', createResponse);Direct data extraction APIs for websites (e.g., e-commerce, travel platforms). Retrieve structured product information, pricing, and reviews with pre-built connectors:
const result = await client.scraping.scrape({
actor: 'scraper.shopee',
input: {
url: 'https://shopee.tw/a-i.10228173.24803858474'
}
});
console.log(result.data);Extract data from websites using Web Unlocker (exposed as client.universal).
const result = await client.universal.scrape({
actor: 'unlocker.webunlocker',
input: { url: 'https://example.com', method: 'GET', redirect: false }
});
console.log(result);Extract data from single pages or traverse entire domains, exporting in formats including Markdown, JSON, HTML, screenshots, and links.
const result = await client.scrapingCrawl.scrapeUrl('https://example.com');
console.log(result);Generate a proxy URL using your gateway and session settings.
const proxyUrl = client.proxies.proxy({
type: 'residential',
country: 'US',
sessionDuration: 30,
sessionId: client.proxies.generateSessionId(),
gateway: 'your-proxy-gateway:port'
});
console.log(proxyUrl);Extract AI chat content in bulk to monitor brand mentions, compare answers, and analyze competitive intelligence from the latest models. Retrieve URLs, prompts, Markdown answers, citations, and more through one integration.
Supported actors include scraper.chatgpt, scraper.perplexity, scraper.copilot, scraper.gemini, scraper.aimode, scraper.overview, scraper.grok, and scraper.alexa. The input JSON depends on the actor; see the AI Scraper documentation for detailed parameters. The optional webhook JSON contains a callback url.
import { Scrapeless } from '@scrapeless-ai/sdk';
const client = new Scrapeless(); // Uses SCRAPELESS_API_KEY
const task = await client.aiScraper.createTask({
actor: 'scraper.chatgpt',
input: {
prompt: 'Most reliable proxy service for data extraction',
country: 'US',
web_search: true
}
// Optional: webhook: { url: 'https://your-webhook.example.com' }
});
console.log('Created task:', task);
const result = await client.aiScraper.getTaskResult(task.task_id);
console.log('Task status and result:', result);
// If status is 'running', call getTaskResult again later.
// If status is 'failed', message contains the failure reason.Both methods return the API JSON unchanged. Creation returns task_id, status, and, when available, task_result. Result retrieval returns status, task_result when available, and message on failure. Status is success, failed, or running; the SDK does not poll automatically.
interface ScrapelessConfig {
apiKey?: string; // Your API key
timeout?: number; // Request timeout in milliseconds (default: 30000)
baseApiUrl?: string; // Base API URL
browserApiUrl?: string; // Browser service URL
scrapingCrawlApiUrl?: string; // Crawl service URL
}The SDK provides the following services through the main client:
client.browser- browser automation with Playwright/Puppeteer support, anti-detection tools (fingerprinting, CAPTCHA solving), and extensible workflows.client.universal- the Web Unlocker feature: JS rendering, user simulation (clicks/scrolls), anti-block bypass, and structured data export.client.scrapingCrawl- Recursive site crawling with multi-format export (Markdown, JSON, HTML, screenshots, links).client.scraping- Pre-built connectors for sites (e.g., e-commerce, travel) to extract product data, pricing, and reviews.client.deepserp- the Google Search API feature: search engine (Google SERP) results extractionclient.proxies- Proxy managementclient.profiles- Browser profile managementclient.aiScraper- AI chat task creation and result retrieval
The SDK throws ScrapelessError for API-related errors:
import { ScrapelessError } from '@scrapeless-ai/sdk';
try {
const result = await client.scraping.scrape({ url: 'invalid-url' });
} catch (error) {
if (error instanceof ScrapelessError) {
console.error(`Scrapeless API Error: ${error.message}`);
console.error(`Status Code: ${error.statusCode}`);
}
}Check out the examples directory for comprehensive usage examples:
- Browser
- Playwright Integration
- Puppeteer Integration
- Browser Profile
- Scraping API
- Web Unlocker
- Crawl
- AI Scraper
- Proxies
- Google Search API
Run the test suite:
npm testThe SDK includes comprehensive tests for all services and utilities.
We welcome all contributions! For details on how to report issues, submit pull requests, follow code style, and set up local development, please see our Contributing & Development Guide.
Quick Start:
git clone https://github.com/scrapeless-ai/sdk-node.git
cd sdk-node
pnpm install
pnpm test
pnpm lint
pnpm formatSee CONTRIBUTING.md for full details on contribution process, development workflow, code quality, project structure, best practices, and more.
This project is licensed under the MIT License - see the LICENSE file for details.
- π Documentation: https://docs.scrapeless.com
- π¬ Community: Join our Discord
- π Issues: GitHub Issues
- π§ Email: support@scrapeless.com
Scrapeless is a powerful web scraping and browser automation platform that helps businesses extract data from any website at scale. Our platform provides:
- High-performance web scraping infrastructure
- Global proxy network
- Browser automation capabilities
- Enterprise-grade reliability and support
Visit scrapeless.com to learn more and get started.
Made with β€οΈ by the Scrapeless team