# Capturing Webpages, including subpages

**URL:** https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897
**Category:** DEVONthink
**Created:** [May 6, 2022, 9:49am UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897 "2022-05-06T09:49:15Z")
**Posts on this page:** 14
**Page:** 2

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 10, 2022, 4:07pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/21 "2022-05-10T16:07:46Z")

</div>

Ok, Sitesucker is doing that stuff and I purchased the pro version. The thing is that my first attempt with DT got me much further than spending an hour with SiteSucker. The only thing missing with DT was that links in the main page need to be linked to the downloaded location. Otherwise DT does the job much easier, in my view.

My scenario is like downloading a Wikipedia page, with a predefined level of depth, for offline reading. I wish DT could do that as well, being so close (and in my view easier than with Sitesucker where far more parameters are available and need to be understood).

---

<div class="post-metadata">

### Author: ![cgrunenberg](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/cgrunenberg/32/7172_2.png) [@cgrunenberg](https://discourse.devontechnologies.com/u/cgrunenberg)
#### Post date: [May 10, 2022, 4:32pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/22 "2022-05-10T16:32:29Z")

</div>

> [@Wolkenhauer](#):
>
> My scenario is like downloading a Wikipedia page, with a predefined level of depth, for offline reading. I wish DT could do that as well, being so close (and in my view easier than with Sitesucker where far more parameters are available and need to be understood).

Could you please send us an example URL plus either an exact description what actually should be imported and how or just send us the Sitesucker result for comparison? Thank you!

---

<div class="post-metadata">

### Author: ![rmschne](https://discourse.devontechnologies.com/letter_avatar_proxy/v4/letter/r/8baadc/32.png) [@rmschne](https://discourse.devontechnologies.com/u/rmschne)
#### Post date: [May 10, 2022, 4:48pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/23 "2022-05-10T16:48:46Z")

</div>

Did you ever check out the “free” curl ? Does not that work?

---

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 10, 2022, 5:21pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/24 "2022-05-10T17:21:26Z")

</div>

I can try but the point of my comment here was that DT has this functionality, I use the web archive every day, I only want one level further with links from that page to have that material offline available. I was therefore hoping that it can be done with DT. All files are downloaded, “all” that is missing is the links in the webpage being redirected to the downloaded files 🙂 I was hoping that this can be done, only that I am too stupid to choose the right option 😉

---

<div class="post-metadata">

### Author: ![chrillek](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/chrillek/32/29340_2.png) [@chrillek](https://discourse.devontechnologies.com/u/chrillek)
#### Post date: [May 10, 2022, 5:30pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/25 "2022-05-10T17:30:33Z")

</div>

Curl does not rewrite URLs, wget does. But the OP doesn’t want to use a CLI, if I understood them correctly.

---

<div class="post-metadata">

### Author: ![rmschne](https://discourse.devontechnologies.com/letter_avatar_proxy/v4/letter/r/8baadc/32.png) [@rmschne](https://discourse.devontechnologies.com/u/rmschne)
#### Post date: [May 10, 2022, 5:33pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/26 "2022-05-10T17:33:10Z")

</div>

either curl or wget will probably work for what is describes as a need. searching for alternatives without trying seems not how i would work the issue. just me i guess.

---

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 10, 2022, 5:44pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/27 "2022-05-10T17:44:36Z")

</div>

Here how it works in DT. The arrow points at the file to which the webpage that I want to download, refers to. The file is downloaded but the links therein do not point to the subpages, that were downloaded nicely with DT:

 ![2022-05-10_19-39-30](https://devontech-discourse.s3.dualstack.us-east-1.amazonaws.com/uploads/original/3X/e/e/ee276e11755b0306f8ebad40aa86b224dc41dab3.png)

I actually could not get this running, that well, with Sitesucker. I am trying to create a simple example from a Wikipedia page, the problem is that most pages there point to so many things that the download is massive. I have to think of a topic where there is little written about 🙂

---

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 10, 2022, 5:51pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/28 "2022-05-10T17:51:50Z")

</div>

Here is an example to test this with Wikipedia:

[https://en.wikipedia.org/wiki/Symplectic\_basis](https://en.wikipedia.org/wiki/Symplectic_basis)

Sitesucker downloads 44 files for a depth of two levels but the nice thing is that in the html file Sympletic\_basis.html all links point to the downloaded files.

 ![simpletic_basis](https://devontech-discourse.s3.dualstack.us-east-1.amazonaws.com/uploads/original/3X/3/8/3886ffc0d01349e4a3308a44420fb3647d01ece2.jpeg)

So strangely, for my actual example DT does well, just short of the redirection of links and Sitesucker does not work. For the Wikipedia example Sitesucker produces the result.

I guess its not an easy task but seems like something that is “just” one level deeper than web archive and thus suiting DT well.

Thanks everyone for commenting!

Olaf

---

<div class="post-metadata">

### Author: ![chrillek](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/chrillek/32/29340_2.png) [@chrillek](https://discourse.devontechnologies.com/u/chrillek)
#### Post date: [May 10, 2022, 6:03pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/29 "2022-05-10T18:03:09Z")

</div>

> [@Wolkenhauer](#):
>
> Thanks everyone for commenting!

Several people have already suggested tools that are _known_ to work. Apparently, you decided to not try them. I for one don’t feel inclined to comment on another tool which I do not even know.

---

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 10, 2022, 6:16pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/30 "2022-05-10T18:16:15Z")

</div>

As I wrote, I purchased Sitesucker, when it was recommended above. I think this qualifies for trying. Not sure what your comment is trying to suggest. I even documented the experiments. So, what did I comment about without knowing or trying?

---

<div class="post-metadata">

### Author: ![chrillek](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/chrillek/32/29340_2.png) [@chrillek](https://discourse.devontechnologies.com/u/chrillek)
#### Post date: [May 10, 2022, 6:21pm UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/31 "2022-05-10T18:21:42Z")

</div>

See below

> [@rmschne](#):
>
> Did you ever check out the “free” curl ? Does not that work?

> [@chrillek](#):
>
> I’d suggest you go for a `wget`, `curl` or such tools.

---

<div class="post-metadata">

### Author: ![cgrunenberg](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/cgrunenberg/32/7172_2.png) [@cgrunenberg](https://discourse.devontechnologies.com/u/cgrunenberg)
#### Post date: [May 11, 2022, 9:01am UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/32 "2022-05-11T09:01:40Z")

</div>

> [@Wolkenhauer](#):
>
> Sitesucker downloads 44 files for a depth of two levels but the nice thing is that in the html file Sympletic\_basis.html all links point to the downloaded files.

I’m not sure which setting you used but a depth of two levels should download many more files actually. Were the links limited to the same host or subdirectory maybe? Did Sitesucker import the resources (e.g. images, stylesheets and scripts) of the pages too?

---

<div class="post-metadata">

### Author: ![Wolkenhauer](https://discourse.devontechnologies.com/user_avatar/discourse.devontechnologies.com/wolkenhauer/32/10482_2.png) [@Wolkenhauer](https://discourse.devontechnologies.com/u/Wolkenhauer)
#### Post date: [May 11, 2022, 9:25am UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/33 "2022-05-11T09:25:51Z")

</div>

Yes, Sitesucker loads down everything, including the images and maths. The pages are recreated as if they are online but the links from the main document point to and open links in the downloaded folder.

My use case is just text, basically a webpage with an article split into several sub-webpages with text, so that the DT web archive functionality works well (for a single page).

---

<div class="post-metadata">

### Author: ![system](https://devontech-discourse.s3.dualstack.us-east-1.amazonaws.com/uploads/original/2X/4/4e6ff70c96be2eb307a274d890bba8d71ba671b3.png) [@system](https://discourse.devontechnologies.com/u/system)
#### Post date: [May 10, 2025, 9:26am UTC](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897/34 "2025-05-10T09:26:26Z")

</div>

This topic was automatically closed 1095 days after the last reply. New replies are no longer allowed.

[Previous page](https://discourse.devontechnologies.com/t/capturing-webpages-including-subpages/70897.md?page=1)
