scrapemate

Scrapemate is a web crawling and scraping framework written in Golang. It is designed to be simple and easy to use, yet powerful enough to handle complex scraping tasks.

Features

Low level API & Easy High Level API
Customizable retry and error handling
Javascript Rendering with ability to control the browser
Screenshots support (when JS rendering is enabled)
Capability to write your own result exporter
Capability to write results in multiple sinks
Default CSV writer
Caching (File/LevelDB/Custom)
Custom job providers (memory provider included)
Headless and Headful support when using JS rendering
Automatic cookie and session handling
Rotating HTTP/HTTPS/SOCKS5 proxy support

JavaScript Rendering

Scrapemate uses Playwright for JavaScript rendering. It requires the Playwright browsers to be installed.

# Install playwright browsers
go run github.com/playwright-community/playwright-go/cmd/playwright install --with-deps chromium

Build and run without any special tags:

go build ./...

Example Usage

The books-to-scrape-simple example demonstrates JavaScript rendering with Playwright:

# Run with Playwright (default)
# First install browsers: go run github.com/playwright-community/playwright-go/cmd/playwright install --with-deps chromium
go run . -js

Installation

go get github.com/gosom/scrapemate

Quickstart

package main

import (
	"context"
	"encoding/csv"
	"fmt"
	"net/http"
	"os"
	"strings"
	"time"

	"github.com/PuerkitoBio/goquery"
	"github.com/gosom/scrapemate"
	"github.com/gosom/scrapemate/adapters/writers/csvwriter"
	"github.com/gosom/scrapemate/scrapemateapp"
)

func main() {
	csvWriter := csvwriter.NewCsvWriter(csv.NewWriter(os.Stdout))

	cfg, err := scrapemateapp.NewConfig(
		[]scrapemate.ResultWriter{csvWriter},
	)
	if err != nil {
		panic(err)
	}
	app, err := scrapemateapp.NewScrapeMateApp(cfg)
	if err != nil {
		panic(err)
	}
	seedJobs := []scrapemate.IJob{
		&SimpleCountryJob{
			Job: scrapemate.Job{
				ID:     "identity",
				Method: http.MethodGet,
				URL:    "https://www.scrapethissite.com/pages/simple/",
				Headers: map[string]string{
					"User-Agent": scrapemate.DefaultUserAgent,
				},
				Timeout:    10 * time.Second,
				MaxRetries: 3,
			},
		},
	}
	err = app.Start(context.Background(), seedJobs...)
	if err != nil && err != scrapemate.ErrorExitSignal {
		panic(err)
	}
}

type SimpleCountryJob struct {
	scrapemate.Job
}

func (j *SimpleCountryJob) Process(ctx context.Context, resp *scrapemate.Response) (any, []scrapemate.IJob, error) {
	doc, ok := resp.Document.(*goquery.Document)
	if !ok {
		return nil, nil, fmt.Errorf("failed to cast response document to goquery document")
	}
	var countries []Country
	doc.Find("div.col-md-4.country").Each(func(i int, s *goquery.Selection) {
		var country Country
		country.Name = strings.TrimSpace(s.Find("h3.country-name").Text())
		country.Capital = strings.TrimSpace(s.Find("div.country-info span.country-capital").Text())
		country.Population = strings.TrimSpace(s.Find("div.country-info span.country-population").Text())
		country.Area = strings.TrimSpace(s.Find("div.country-info span.country-area").Text())
		countries = append(countries, country)
	})
	return countries, nil, nil
}

type Country struct {
	Name       string
	Capital    string
	Population string
	Area       string
}

func (c Country) CsvHeaders() []string {
	return []string{"Name", "Capital", "Population", "Area"}
}

func (c Country) CsvRow() []string {
	return []string{c.Name, c.Capital, c.Population, c.Area}
}

go mod tidy
go run main.go 1>countries.csv

(hit CTRL-C to exit)

Migrating from v0.9.x to v1.0.0

Version 1.0.0 introduces a BrowserPage interface abstraction for browser automation. This is a breaking change for users who use JavaScript rendering with BrowserActions.

Update `BrowserActions` signature

// Before (v0.9.x)
func (j *MyJob) BrowserActions(ctx context.Context, page playwright.Page) scrapemate.Response {
    page.Goto("https://example.com", playwright.PageGotoOptions{
        WaitUntil: playwright.WaitUntilStateNetworkidle,
    })
    html, _ := page.Content()
    return scrapemate.Response{Body: []byte(html)}
}

// After (v1.0.0)
func (j *MyJob) BrowserActions(ctx context.Context, page scrapemate.BrowserPage) scrapemate.Response {
    resp, err := page.Goto("https://example.com", scrapemate.WaitUntilNetworkIdle)
    if err != nil {
        return scrapemate.Response{Error: err}
    }
    return scrapemate.Response{
        Body:       resp.Body,
        StatusCode: resp.StatusCode,
    }
}

Accessing the underlying browser page

If you need browser-specific features, use Unwrap():

// For Playwright
pwPage := page.Unwrap().(playwright.Page)

See CHANGELOG.md for the full list of changes.

Documentation

You can find more documentation here

For the High Level API see this example.

Contributing

Contributions are welcome.

Licence

Scrapemate is licensed under the MIT License. See LICENCE file

Name		Name	Last commit message	Last commit date
Latest commit History 83 Commits
.github/workflows		.github/workflows
adapters		adapters
cmd		cmd
docs		docs
examples		examples
mock		mock
scrapemateapp		scrapemateapp
.gitignore		.gitignore
.golangci.yaml		.golangci.yaml
.pre-commit-config.yaml		.pre-commit-config.yaml
CHANGELOG.md		CHANGELOG.md
LICENSE		LICENSE
Makefile		Makefile
README.md		README.md
browser.go		browser.go
constants.go		constants.go
context.go		context.go
errors.go		errors.go
go.mod		go.mod
go.sum		go.sum
job.go		job.go
lint.go		lint.go
proxy.go		proxy.go
proxy_test.go		proxy_test.go
response.go		response.go
result.go		result.go
scrapemate.go		scrapemate.go
scrapemate_test.go		scrapemate_test.go
services.go		services.go

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

scrapemate

Features

JavaScript Rendering

Example Usage

Installation

Quickstart

Migrating from v0.9.x to v1.0.0

Update `BrowserActions` signature

Accessing the underlying browser page

Documentation

Contributing

Licence

About

Uh oh!

Releases 19

Packages

Uh oh!

Uh oh!

Contributors

Uh oh!

Languages

Folders and files

Latest commit

History

Repository files navigation

scrapemate

Features

JavaScript Rendering

Example Usage

Installation

Quickstart

Migrating from v0.9.x to v1.0.0

Update BrowserActions signature

Accessing the underlying browser page

Documentation

Contributing

Licence

About

Topics

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases 19

Packages 0

Uh oh!

Uh oh!

Contributors

Uh oh!

Languages

Update `BrowserActions` signature

Packages