Updated curl, exiftool, Java, jq, Python, Tesseract-OCR and even more important, finally switched from sequential calls with tika-app for file content extraction to using tika-pipes! Result is an enormous performance boost for content extraction because object stays alive and multiple clients and emitters are used in parallel. Processing time for the default test data set went down from 350+ to 8 seconds! 😱
INFO [pool-2-thread-9] 01:23:35,996 org.apache.tika.pipes.PipesClient pipesClientId=7: starting process INFO [pool-2-thread-2] 01:23:35,996 org.apache.tika.pipes.PipesClient pipesClientId=0: starting process ... INFO [pool-2-thread-6] 01:23:36,001 org.apache.tika.pipes.PipesClient pipesClientId=3: starting process INFO [main] 01:23:44,087 org.apache.tika.async.cli.TikaAsyncCLI Successfully finished processing 45 files in 8219 ms
Next
Further testing of pipes exploring optimal parameters settings and running on larger datasets.
Er zijn projecten die op papier klein ogen, maar een kijkje onder de motorkap (en het nodige denkwerk om het concept te snappen) laat zien hoe je als developer met de juiste tools veel tijd en resources bespaart. NativeLink is er zo eentje en als (voor mij) eerste kennismaking heb ik er Simpal mee gemaakt. Aan de buitenkant een simpele rekenmachine (ToyCalc), maar ook een showcase van hoe je het buildproces slimmer en sneller maakt.
Met NativeLink van Trace Machina dus. De belofte vooraf: na een paar keer bouwen met NativeLink voelt lokaal compilen ineens alsof je terug bent in de tijd van dial-up internet, oef…
De façade: ToyCalc
ToyCalc zelf is niets meer dan een simpele calculator in Python. Het draait hier echter niet om het rekenen maar om de setup. De manier waarop dependencies, distributies en builds zijn geregeld, dat is waar de magie zit.
Simpal gebruikt Bazel, een serieus stuk gereedschap dat bekendstaat om z’n deterministische builds en herbruikbare artefacten. En precies daar haakt NativeLink slim op in.
De “cheat code”: NativeLink in je build
Kijk naar onderstaand stukje van build.sh. Hier zie je meteen dat NativeLink letterlijk een schakelaar is voor full remote execution of gewoon caching met een environment variable:
if [ "$USE_REMOTE_EXEC" = "1" ]; then
echo "[INFO] Using full remote execution via NativeLink (--config=remote-exec)"
BAZEL_FLAGS="--config=remote-exec"
else
echo "[INFO] Using local execution with remote cache (default)"
fi
bazel build //:build_pyinstaller $BAZEL_FLAGS --verbose_failures
Oftewel:
Zet USE_REMOTE_EXEC=1 en je build gaat de cloud in, draait daar snel z’n rondje en komt terug met een keurig pakketje.
Laat je ‘m uit, dan gebruik je NativeLink “gewoon” als slimme cache.
Je hoeft verder niets te doen dan wel/niet kiezen voor de cloud build en ineens gaat het veeeel sneller.
NativeLink via Nix: een helper met smaak
De code is geschreven voor macOS waarbij NativeLink wordt geïntegreerd zodat je het via een shell-functie kunt gebruiken. Dus kan zijn dat je voor ander OS wat moet aanpassen.
NATIVELINK_HELPER='nativelink() {
if [[ "$1" == "--help" || "$1" == "--version" ]]; then
nix run github:TraceMachina/nativelink \
--extra-experimental-features nix-command \
--extra-experimental-features flakes -- "$@"
else
nix run github:TraceMachina/nativelink \
--extra-experimental-features nix-command \
--extra-experimental-features flakes \
"$HOME/GitHub/taatuut/simpal/basic_cas.json5" "$@"
fi
}'
Dat betekent dat je (na één keer je terminal opnieuw starten) je gewoon nativelink build kunt tikken zonder de extra flags.
Waarom is dit handig
Voorkomt “het werkt -alleen- op mijn machine” en versnelt het build proces.
Met NativeLink worden builds consistent, cachebaar en reproduceerbaar waarbij collega’s en de CI-server dezelfde artefacten gebruiken.
De snelheid krijg je ‘kado’ door gebruik maken van cloud infrastructuur en bovenal het feit dat NativeLink alleen dat opnieuw doet waar veranderingen in de code aan ten grondslag liggen. In mijn simpele calculator projectje zie je vooral verschil tussen eerste build en navolgende. Zonder veranderingen in de build (ja waarom dan builden, maar even als praktische verificatie van de theorie) gaat dan dus razendsnel, en ook bij wijzigingen is het sneller vanwege het hierboven genoemde: alleen doen wat nodig is. Echte winst ga je daar natuurlijk vooral mee halen in grote(re) projecten die vaak gebuild, getest en gereleased worden.
Wat heb ik nu weer geleerd
Je hebt niet per se een megaproject nodig hebt om met NativeLink te werken. Een kleine repo als ToyCalc is genoeg om het nut te laten zien. En als dit werkt voor deze demo, werkt het nog beter voor een enterprise-monorepo met 600 microservices.
Is remote de toekomst voor alle builds? Dat weet ik niet, maar voor enterprise projects met meerdere participanten is het zeker een valide optie zo te zien aan de bedrijven op de NativeLink website die er al mee bezig zijn.
NativeLink verandert hiermee een spelregel van software bouwen waardoor traaaage en inconsistente builds niet langer nodig zijn.
About seven years after creating the first blend, following a too early announcement in 2021 it now really happened on April 27th, 2024. More details like label and availability to be released later, let’s start sharing the tasting notes.
Nose: at first very light, some vanilla hiding some caramel under a blanket of fresh distillate. Further away some stranger ground, hard to detect the terroir. Mouth: starts quite spiritual, then a short blast of red fruits coming in and eventually back to more delicate and volatile hard to grap ethereal tones. A ticking sensation lasts for a while. If you can resist swallowing directly, a hint of berries (vlierbessen) appears. After a while in the glass, the bouquet opens up and becomes richer with some subtle floral notes and a bit of herbs and wood. It does not last too long so time for another sip.
I will refrain from using That Meme again, although it would actually have been a nice companion for this post. However, it is 2024 and time for Something New, so no meme and instead of that sprinkling Random Uppercase throughout this post, let’s see if we can make a trend out of that (but that is not the point).
Nope, first some Real News (oops, last time I swear) on changes in the ControlOne Agent (these capitals where there already before 2024). Here are the tech updates:
Upped Java to openjdk-21.0.1
Upped Python to PyPy 3.10 v7.3
Upped Tika to 2.9.1
And started first explorations with tika-pipes! It would be too harsh to say that the current way Tika is used with Python and tika-app is not good enough because it serves us well and delivers the required results. But it felt a bit brittle and with recent updates I saw some unexplainable behaviour so probably the right time to make a move, also because I could not find a current and more extensive Python module alternative for Tika.
Even iets heel anders dan mijn gebruikelijke stokpaardjes als geospatial, whisky en data in al zijn verschijningsvormen. Al decennialang ben ik gefascineerd door films als Convoy en Breaker! Breaker! (trouwens eigenlijk alles van Chuck Norris zeg ik er eerlijkheidshalve bij).
Omdat ik merk dat de kennis over deze films bij de huidige generatie stilaan verloren begint te gaan, heb ik de Rubber Duck maar eens bij de spreekwoordelijke veren gevat en een korte studie geschreven over de verschillen tussen deze twee films waarbij rekening wordt gehouden hoe zij beiden in de tijdgeest van de 70er jaren passen. Daarbij betrek ik ook de invloed van de regisseurs en hoe hun achtergrond in eventuele eerdere en latere films van hun hand zich ook toont in respectievelijk Breaker! Breaker! en Convoy. Overigens dacht ik altijd dat Breaker! Breaker! na Convoy uitkwam en ik was zelfs een tijdje bang dat Chuck een goedkope rip-off wilde maken, letterlijk & figuurlijk meeliftend op de populariteit van the truckers film genre. Gelukkig bleek niks minder waar en was Breaker! Breaker! zelfs eerder uit dan Convoy! Ik had het kunnen weten, Chuck faalt nooit (als filmheld hè… maar wat de…? https://chucknorrisfacts.net/ is niet langer online? Drat, drat & double drat!).
Tijdens dit korte onderzoek (want ik was eigenlijk met iets anders bezig, daarover later meer) moest ik natuurlijk ook luisteren naar de bijbehorende truckers muziek. Om mezelf niet te veel af te laten leiden, heb ik in eerste instantie gekozen vooral Nederlandse liedjes uitgebracht na de releases van Convoy en Breaker! Breaker!
Met gepaste trots op het Nederlandse truckerslied presenteer ik u een fijne selectie:
B.B. Band – Stille Willie
Tina Trucker – Lady Trucker is mijn naam
Kenteken Onbekend – Hou ‘m tussen de strepen
En eigenlijk valt de volgende buiten de strenge selectiecriteria maar omdat de naam van zowel artiest als liedje zo koel zijn, zet ik hem er toch bij:
Don Mercedes – Rocky
(hiervan vind ik trouwens dat het loopje dat her en der opduikt, erg lijkt op de bas uit Seasons van Future Islands, vooruit hier de link dan kan je het zelf checken https://www.youtube.com/watch?v=upPl9mZW_zw)
Natuurlijk kon ik niet anders dan Convoy van C.W. McCallhttps://www.youtube.com/watch?v=wwaygKjs2fI meenemen in de luistersessies en al bij de eerste herbeluistering na jaren viel me meteen op dat het refrein eigenlijk in een soort van South Park stijl avant la lettre is.
‘Cause we got a little convoy Rockin’ through the night. Yeah, we got a little convoy, Ain’t she a beautiful sight? Come on and join our convoy Ain’t nothin’ gonna get in our way. We gonna roll this truckin’ convoy ‘Cross the U-S-A. Convoy!
Dit inspireerde me om ChatGPT DALL·E 3 te vragen een aantal bijpassende afbeeldingen te maken die ik in de studie op redelijk willekeurige plekken toegevoegd heb. Zelf vind ik daarin vooral de vrijzinnige mix van de Southpark karakters en het niet geheel ontbreken van krachttermen interessant. En opnieuw onder de indruk van wat ChatGPT DALL·E 3 kan produceren. Nou daar komt ‘ie en vooral benieuwd of er nog aanvullingen zijn met betrekking tot eventuele inzichten die ik heb gemist.
“Breaker! Breaker!” (1977) en “Convoy” (1978) zijn twee films uit de late jaren ’70 die de truckerscultuur en het CB-radiofenomeen (Citizens’ Band radio oftewel 27MC bakkies) in de Verenigde Staten belichten. Ze passen in de tijdgeest van de jaren ’70, een periode waarin de truckercultuur en de CB-radio een hoogtepunt bereikten in populariteit. Deze films bieden interessante inzichten in de culturele en cinematografische trends van die tijd, alsook in de carrières van hun regisseurs, Don Hulette en Sam Peckinpah (ook zo’n ontzettende aanrader, alle films van Sam, misschien iets voor een volgende off-topic post).
Breaker! Breaker!
Geregisseerd door Don Hulette, vertelt “Breaker! Breaker!” het verhaal van J.D., een trucker gespeeld door Chuck Norris, die tegen corruptie vecht in een klein stadje. De film is een actiefilm en vertegenwoordigt een van de vroegere rollen van Norris, die toen nog aan het begin van zijn filmcarrière stond. Norris zelf beschreef de film als een “down-home soort film”. Merk op dat de film in slechts elf dagen werd opgenomen met een beperkt budget.
De regie van Hulette in “Breaker! Breaker!” kan worden gezien als een voorbeeld van onafhankelijke filmmaken in de jaren ’70. Zijn aanpak is rechttoe rechtaan, zonder veel flair of complexiteit, wat past bij de beperkte middelen en snelle productietijd (denk in dit verband ook aan de Dirty Harry en Death Wish films). Het was een van de slechts twee films die Hulette regisseerde, naast “A Great Ride”.
Convoy
“Convoy”, geregisseerd door Sam Peckinpah, is gebaseerd op het gelijknamige country-westernlied van C.W. McCall en speelt zich af tegen de achtergrond van de truckers- en CB-radio-cultuur. Met toen al bekende sterren als Kris Kristofferson, Ali MacGraw en Ernest Borgnine, vertelt de film het verhaal van truckers die zich verzetten tegen corrupte politieagenten (‘bears’) en autoriteiten.
Peckinpah, bekend om zijn eerdere werken zoals “The Wild Bunch” en “Straw Dogs”, had een turbulente carrière gekenmerkt door zijn strijd met alcoholisme en drugsverslaving. “Convoy” was een poging om zijn carrière weer op de rails te krijgen na een reeks commercieel teleurstellende films. Zijn stijl in “Convoy” weerspiegelt een mengeling van zijn kenmerkende gewelddadige esthetiek en een meer mainstream, commercieel georiënteerde benadering. Ondanks gemengde kritieken werd de film de meest commercieel succesvolle van Peckinpah’s carrière.
Vergelijking en Tijdgeest
Beide films weerspiegelen de fascinatie van de jaren ’70 voor de truckercultuur en CB-radio’s, een fenomeen dat ook zichtbaar was in andere media van die tijd maar bieden verschillende benaderingen van dit thema. “Breaker! Breaker!” leunt meer naar de kant van een onafhankelijke actiefilm, terwijl “Convoy” een grotere, meer dramatische productie is met een bekendere regisseur en cast.
De invloed van de regisseurs is duidelijk in beide films. Hulette’s werk in “Breaker! Breaker!” is eenvoudig en ongecompliceerd, passend bij de beperkte middelen en de vroege carrièrefase van Norris. Aan de andere kant laat Peckinpah’s “Convoy” zijn worstelingen en ambities zien om zowel commercieel succes als artistieke erkenning te bereiken, ondanks de uitdagingen van zijn persoonlijke leven en carrière.
In de context van de jaren ’70 vertegenwoordigen deze films de brede aantrekkingskracht van de truckerscultuur en een romantische kijk op de vrijheid van de open weg, thema’s die resoneren met de bredere Amerikaanse cultuur van dat decennium.
Nou dat was hem. Roept meteen de vraag op hoe de Amerikaanse (truckers)cultuur de Nederlandse beïnvloed heeft. En wat de mooiste truck was: Kentucky, Peterbilt, International… In Nederland in ieder geval de Scania met neus destijds. Oh ja, je had ook nog van die 27MC kaartjes van truckers, ook nog een tijdje verzameld. Voer voor vervolgen maar niet nu.
~Breaker one-nine, this here’s the Rubber Duck!~
Met dank aan onder andere Wikipedia en ongebrande studentenhaver (met uitzondering van de hazelnoten).
The DataFitness ControlOne Agent is already doing its job of collecting vast amounts of data for about three years. The first ‘official’ commit to the Dataether DataFitness repository was on the day after my birthday in November 2020, and before that there were some initial versions stored in my private repos.
We have been tracking data developments over time at one company for close to two years know, really disclosing interesting insights and monitoring information fit for governance and general data management purposes.
Also, the collected data can be used in many different ways, including full-text fuzzy search and vector search (or ‘hybrid search’ as the combination is called nowadays), delivering additional information that can be fed back into the company for direct action. We even did a first successful test creating executable scripts for mass data movements based on customer driven rules and decision criteria (yes, the scripts were actually executed).
Development on the Agent is still ongoing. A small list of recent changes:
Updated Tika to version 2.8.0, this should give more detail on geospatial data. Some fixes to deal with the extracted data in a better way.
One of my geospatial stokpaardjes is nagging about (too many) decimals in coordinates. Someone else has been kind enough to wrap that message in a funny & understandable way, see https://xkcd.com/2170/
Recently I was asked about options to do ‘find nearest’ geospatial calculations for a set with lots of polygon features, and whether something like GeoSPARQL, a (spatial) database or something else like Python and GeoPandas would be a good fit. The exact question was finding the nearest top 3 Natura2000 areas for all Natura2000 areas. This came up in some of the project work Sensing Clues is doing in cooperation with their field partners.
My advise after some discussion was not to do geospatial calculations, at least not more often than necessary, meaning don’t do them runtime but run the operation upfront and store the results as properties with the data (e.g. additional triples in a graph database, some optional fields in a document oriented data platform, or just a good-old lookup table in a relational database).
I started with the Natura2000 data in Esri Shapefile format, wrote some Python code to load and process it with GeoPandas and ran the job. No optimisations, not split to run stuff in parallel or whatever. After about 2-3 days only 4000 areas were done (single process execution on my 2015 MacBook Pro), and I did not want to wait anymore so changed the approach and switched to PostgreSQL/PostGIS.
Decided to use the Natura2000 Geopackage instead of the Shapefile because running find nearest after Shapefile import in PostGIS seemed to be slower than using the Geopackage. Disclaimer: no idea yet why imported Shapefile is slower, no deep investigation done, maybe number of decimals after import, or difference in complexity of the (multi) polygon features, or the way either data from Shapefile or Geopackage are spatially indexed by PostGIS. Of course ogr2ogr helped with loading into PostGIS.
Also started looping over the Natura2000 features instead of one time join all, wrote a procedure for this (prefer procedure over function because in procedure you can commit in between and track progress easier). Some features were still slow to process because they are very big and complex, and/or so are their neighbors. Some Natura2000 areas can span complete rivers and waterway systems covering tens to hundreds of square kilometers.
With this approach more than 10.000 features were done in less than a day, still with the slower ones consuming a fair amount of time. In less than two days the whole set of 27020 areas finished. A quick examination of the results revealed that the results of PostgreSQL/PostGIS and Python/Geopandas find nearest operation are the same most of the tim e. Again disclaimer that I did not check where and why they differ. But I’m sure PostGIS is ok 🙂
One optimisation that might make calculations way simpler and faster, is doing rubberband or envelope on these areas with ST_ConvexHull or ST_OrientedEnvelope, then do nearest on these ‘bounding polygons’.
In this case doing geospatial calculations only once and then add the results as properties or in a lookup table works because the set of Natura2000 areas does not change often. So redoing the whole set when needed is way cheaper then running nearest operations every time where doing analytics for specific areas. And in real life there are many more of these cases where preparing the data by creating smart lookups works fine so it can be a time and resource saver to consider.
The full description and code used can be found at https://github.com/taatuut/NATURA9000 It also includes some example code to visualise the results. Of course quick&dirty code but it should work for you too in case you want to test it.
If you find this interesting and have ideas how to improve this, or want to help Sensing Clues anyway with the work they do then let me know, and check out https://sensingclues.org/ anyway. Also looking for volunteers with GeoServer skills!