we are amazed with your software.
Could you please help with :
We finish downloading a website like www.waqfeya.com for example, with url substitue and additional=keepprimary ...
we browse it like a shame however
when we try to download missing files we get some of the files already downloaded
when exporting thoses files don't get exported or link translation is bad (no image, no style in some files).
I will mail you with the the project properties
thank you in advance
Thank you!
Best regards,
Oleg Chernavin
MP Staff
XXX/book.php?bid=7813
resalty.XXX/index.php
resalty.XXX/index.php/category-15
This problem persists with other websites
I think it's, maybe related to how big the file is ? I checked and unchecked supress website error but no luck.
And another little problem that got to my nerve is how to add / in the end of url if it has no extension. For example a/am became a/am/ and a/am.exe don't get modified
thank you.
I just happen to pass by and found a new version of oee. I will try it to found if it resolve my problems. Thank you for the hard work.
Oleg.
for this :
>And another little problem that got to my nerve is how to add / in the end of url
>if it has no extension. For example a/am became a/am/ and a/am.exe don't get modified
I found a workaround for it. Thank you however.
http://www.metaproducts.com/download/betas/OEP3908.zip
Oleg.
Oleg.
The same problem persits. I tried to download some files of another website and stuck with same problem, http://alhazme.net/ and a lot of other ones ...
However the parser is a little quick than before.
Oleg.
I will try it with another pc to see changes, and will contact you.
Oleg.
For alhazme the problem is solved. All the website is downloaded, than with ctrl+f5 get only the links with error and one that was downloaded before. After the second download of the only one link remaining no link remain. Another ctrl+F5 to test and no link at all. Thank you very match.
However for the first website the problem remain almost the same :
Some link that where redownloaded every time, now are not redownloaded
Other ones like these are redownloaded every time
http://resalty.waqfeya.com/index.php from http://resalty.waqfeya.com/
and so on.
http://www.waqfeya.com/book.php?bid=1043 from http://www.waqfeya.com/category.php?cid=87
Oleg.
Another sub problem that I were about to tell you when we finish with the problem of missing files download
Is the pdf files in archive.org are not all well downloaded.
I almost downloaded 100 gb of files 3 times with no success
I remarqued three things in this project
* 302 moved files (archive.org) downloaded but not found by oee nor are browsable for the most of them
* that can be due to the same problem
* maybe there is some problem with my url substitute or filters.
The website dedewnet.com is downloaded well with oee last version.
But with the last oep that you gave me there are still missing files to download each time (no substitute nor filters)
only for "www." and ".net" to ".com"
And I mailed you with the project settings for waqfeya.com
Oleg.
Oleg, are you ok ?
Sorry for distrubing you ! but is there anything new. If you didn't receive the settings I will paste them directly here in post.
Waiting for your response.
Thank you!
Oleg.
I did resend you the settings a couple of days in an attached file so you don't get errors. Did you get them ?
I'm waiting for you response and eventual solution.
Oleg.
Is there nything new on my problem dear Oleg ? I did resend you the right settings.
Thank you.
OK. I downloaded the Project partially. It browses well. Downloading using Ctrl+F5 doesn't try to get existing files. What should I look for?
Oleg.
Maybe it's minor problem with my installation because i changed a lot betwenn Oee and oep beta.
I will try to reinstall it or use the pc of another user to see changes.
I will try to post after the week-end.
I tried it in other pcs and the same problem remains in some pages so I changed the url substitutes in the remaining ones :
category-1 (category-1 is a file)
category-1/thesis (category-1 is a folder)
I changed the first to category1 for example (impossible to have a folder and a file with the same name)
For the other links I didn't see the problem's facts ???
See you soon
Good news, the problem is solved it's due partially to the pc used in.
Thank you very match for you time.
Oleg.
Sorry for not answering for a long period.
Oleg, I think it's caused by one/all of the following :
1. temp files not purged well
2. problem with url like this one xxx/cat/ and xxx/cat (file and directory)
2. maybe the use of two version at the same time to download the same project (testing time)
The project just got donwloaded well with some changements after reading some topics.
But I got these minor problems.
1. The OEE window hangs when parsing at some stage
the memory is not at the limit (56 Mb used in average)
and all the files have been downloaded by oee
2. Some errors in the parser (don't know yet if this affects the project or not)
ERROR - ... - Error reading from file: C:\DOCUME~1\XXX\LOCALS~1\Temp\htt67.tmp
Error code=00000000 Referer=http://ia600709.us.archive.org/34/items/waqsfmkn_3/
I assume that the parser fails to parse pdf files to get links from their bookmarks
3. Some times I get Error Out of memory
Is it possible to pause the parser or to skip parsing files by size not by name
because I want to parse pdf files but not bigger once if this affects the parser
P.S. :
If you think it's better to make a new topic just tell me
I'm not sure if it's virus related but no software was interfering, I am sure.
some of the errors
Parser error 2: Out of memory URL: http://ia600402.us.archive.org/9/items/waq98398/98398.pdf
Error reading from file: X:\ia700306.us.archive.org\7\items\41455waq\41455.pdf.primary Error code=00000000 Referer=http://ia700306.us.archive.org/7/items/41455waq/
I have to split them into smaller parts to keep memory usage low. Thank you for the tests!
Oleg.
Never mind it.
Thank you too.
Oleg.
In the project's settings I check "download only missing files" and check the two options below it.
Ctrl+F5 and F9 and after some hours of parsing I get links like those :
http://archive.org/details/xx/
and
http://archive.org/XX/items/xx/
and when downloaded it says 304 not modified
It's strange, is not it ?
Files been queued and downloaded when they do exist already.
And please explain me the difference between URL and HTML Text in URL substitute.
URL and Text. The first changes URLs before they get queued. The last changes text in downloaded HTML files.
Oleg.
Yes for the first one, and get all possible subdirectory is checked also.
Oleg.
I don't understand what you said.
Oleg.
For the moment I don't have a small project.
I will redownload mine (a very big project) and reproduce the issue another time.
Yesterday I export the project and get alot of previous url substitute (not the last ones) ...
So now I am not sure that the problem comes from oee.
In fact, I change the url substitute a lot and don't redownload the project,I just download the missing files.
However Keepprimary is always used.
One of the things causing the problem is checking index downloaded links
when unchecked a lot of files (already downloaded) are not queued and oee crashes less.
Maybe because the index didn't complete adding the file or the index got corrupted at a certain point.
And when I index with "optimize the search index" I don't know when it will end. Is there a message ?
Thank you!
Oleg.