RESEARCH ARTICLES
Manner of death, causes of death and autopsies in infants, children and adolescents: An overview from a German metropolis 2002–2012
Nov 2021 – International Journal of Legal Medicine (Springer Nature)
DOI: 10.1007/s00194-022-00568-y
Fractures and skin lesions in pediatric abusive head trauma: a forensic multi-center study
Dec 2021 – International Journal of Legal Medicine (Springer Nature)
DOI: 10.1007/s00414-021-02751-4
Abusive head trauma in court: a multi-center study on criminal proceedings in Germany
Oct 2020 – International Journal of Legal Medicine (Springer Nature)
DOI: 10.1007/s00414-020-02435-5
Post-mortem estimation of gestational age and maturation of new-borns by CT examination of clavicle length, femoral length and femoral bone nuclei
Jun 2020 – Forensic Science International [ACKNOWLEDGED SUPPORT]
DOI: 10.1016/j.forsciint.2020.110391
Pure Functions in C: A Small Keyword for Automatic Parallelization
May 2020 – International Journal of Parallel Programming
DOI: 10.1007/s10766-020-00660-4
Extending PluTo for Multiple Devices by Integrating OpenACC
March 2018 – 26th Euromicro International Conference on Parallel, Distributed and Network-based Processing (PDP)
DOI: 10.1109/PDP2018.2018.00049
Pure Functions in C: A Small Keyword for Automatic Parallelization
September 2017 – IEEE International Conference on Cluster Computing (CLUSTER)
DOI: 10.1109/CLUSTER.2017.32
Energy-Efficiency and Performance Comparison of Aerosol Optical Depth Retrieval on Distributed Embedded SoC Architectures
July 2017 – In book: Scientific Computing and Algorithms in Industrial Simulations (pp.341-358)
DOI: 10.1007/978-3-319-62458-7_17
VarySched: A Framework for Variable Scheduling in Heterogeneous Environments
September 2016 – IEEE International Conference on Cluster Computing (CLUSTER)
DOI: 10.1109/CLUSTER.2016.19
An efficient geosciences workflow on multi-core processors and GPUs: a case study for aerosol optical depth retrieval from MODIS satellite data
February 2016 – International Journal of Digital Earth 9(8):1-18
DOI: 10.1080/17538947.2015.1130087
Comparison of Acceleration Techniques for Selected Low-Level Bioinformatics Operations + Supplementary Material
February 2016 – Frontiers in Genetics 7
DOI: 10.3389/fgene.2016.00005
Impact of the Scheduling Strategy in Heterogeneous Systems That Provide Co-Scheduling
January 2016 – 1st COSH Workshop on Co-Scheduling of HPC Applications
DOI: 10.14459/2016md1286954
Multicore Processors and Graphics Processing Unit Accelerators for Parallel Retrieval of Aerosol Optical Depth From Satellite Data: Implementation, Performance, and Energy Efficiency
June 2015 – IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 8(5):2306-2317
DOI: 10.1109/JSTARS.2015.2438893
Hardware-Aware Automatic Code-Transformation to Support Compilers in Exploiting the Multi-Level Parallel Potential of Modern CPUs
February 2015 – COSMIC ’15 Proceedings of the 2015 International Workshop on Code Optimisation for Multi and Many Cores
DOI: 10.1145/2723772.2723776
Facilitate SIMD-Code-Generation in the Polyhedral Model by Hardware-aware Automatic Code-Transformation
January 2013 – IMPACT 2013 Volume: 45
DOI: 10.13140/2.1.5066.3368
PATENTS
Feld et al.
Method and computer program for determining a placement of at least one circuit for a reconfigurable logic device
Verfahren und Computerprogramm zur Bestimmung einer Positionierung von mindestens einer Schaltung für eine rekonfigurierbare logische Vorrichtung
Embodiments relate to a method and computer program for determining a placement of at least one circuit for a reconfigurable logic device. The method comprises obtaining (110) information related to the at least one circuit. The at least one circuit comprises a plurality of blocks and a plurality of connections between the plurality of blocks. The plurality of blocks comprise a plurality of logic blocks. The method further comprises calculating (120) a circuit graph based on the information related to the at least one circuit. The circuit graph comprises a plurality of nodes and a plurality of edges. The plurality of nodes represent at least a subset of the plurality of blocks of the at least one circuit and wherein the plurality of edges represent at least a subset of the plurality of connections between the plurality of blocks of the at least one circuit. The method further comprises determining (130) a force-directed layout of the circuit graph. The force-directed layout is based on attractive forces based on the plurality of connections between the plurality of blocks and based on repulsive forces between the plurality of blocks. The method further comprises determining (140) a placement of the plurality of logic blocks onto a plurality of available logic cells of the reconfigurable logic device based on the force-directed layout of the circuit graph.
LISTING IN THE EUROPEAN PATENT OFFICE
Application number: EP20160203521 20161212
Priority number(s): EP20160203521 20161212
Also published as:
CN108228972 (B)
CN108228972 (A)
US10460062 (B2)
US2018165400 (A1)
THESIS‘
FieldPlacer – A flexible, fast and unconstrained force-directed placement method for heterogeneous reconfigurable logic architectures. Dissertation, Universität zu Köln.
The field of placement methods for components of integrated circuits, especially in the domain of reconfigurable chip architectures, is mainly dominated by a handful of concepts. While some of these are easy to apply but difficult to adapt to new situations, others are more flexible but rather complex to realize. This work presents the FieldPlacer framework, a flexible, fast and unconstrained force-directed placement method for heterogeneous reconfigurable logic architectures, in particular for the ever important heterogeneous FPGAs. In contrast to many other force-directed placers, this approach is called ‘unconstrained’ as it does not require a priori fixed logic elements in order to calculate a force equilibrium as the solution to a system of equations. Instead, it is based on a free spring embedder simulation of a graph representation which includes all logic block types of a design simultaneously. The FieldPlacer framework offers a huge amount of flexibility in applying different distance norms (e. g., the Manhattan distance) for the force-directed layout and aims at creating adapted layouts for various objective functions, e. g., highest performance or improved routability. Depending on the individual situation, a runtime-quality trade-off can be considered to either produce a decent placement in a very short time or to generate an exceptionally good placement, which takes longer. An extensive comparison with the latest simulated annealing placement method from the well-known Versatile Place and Route (VPR) framework shows that the FieldPlacer approach can create placements of comparable quality much faster than VPR or, alternatively, generate better placements in the same time. The flexibility in defining arbitrary objective functions and the intuitive adaptability of the method, which, among others, includes different concepts from the field of graph drawing, should facilitate further developments with this framework, e. g., for new upcoming optimization targets like the energy consumption of an implemented design.
DOWNLOAD DIPLOMA THESIS
Effiziente Vektorisierung durch semi-automatisierte Code-Optimierung im Polyedermodell
In den vergangenen zwei Jahrzehnten haben sich Vektoreinheiten als Beschleuniger in CPUs auch im Bereich der PCs etabliert. Allerdings sind selbst moderne Compiler nicht generell in der Lage, dieses Potenzial auszuschöpfen, wenn der Programmierer die entsprechenden Codes nicht explizit für den jeweiligen Beschleuniger schreibt. Transformationen des Codes, die für eine Nutzung der Vektoreinheiten nötig wären, können von vielen aktuellen Compilern nicht oder nicht immer realisiert werden, da die hierzu nötigen mathematischen Operationen und Analysen nicht implementiert oder noch gar nicht entwickelt sind. In dieser Arbeit wurden Methoden zur Transformation von Quellcodes entwickelt und implementiert, die auf Gomory-Cut-Lösungsverfahren der ganzzahligen Optimierung und Konzepten der Lineare n Algebra beruhen und als Vorstufe einer End-Kompilierung (durch Compiler wie den GCC oder ICC) eingesetzt werden können. Der Fokus lag auf der Entwicklung von Codeoptimierungen für eine Vektorisierung von Schleifen in C-Codes durch ein semi-automatisches Framework. Ziel der Transformationen ist eine bessere Nutzung von SIMD-Architekturen, z. B. den SSE-Einheiten aktueller x86-Prozessoren, die Prinzipien lassen sich aber auch auf andere Hardware-Architekturen mit vektorbasierten Befehlssätzen übertragen. Neben Transformationen zur Ermöglichung der Vektorisierung steht die automatische Anpassung an die vorliegende Hardware im Fokus der Betrachtung um eine für die Speichernutzung optimale Transformation durchzuführen. Durch die entwickelten Methoden können erhebliche Speedups erreicht werd en, die durch die Speicherzugriffsoptimierung z.T. deutlich über dem durch die parallelen Berechnungen möglichen Speedup liegen.
DOWNLOAD DISSERTATION
Supervision…
…of numerous final theses (Bachelor and Master) and commitment to the promotion of young talent, e.g. through the nationwide Girls-Day
Appointed member…
…of the program committee High-Performance Computing in Remote Sensing, Sep 2015
Several years reviewer for the IEEE…
Journal of Selected Topics in Applied Earth Observations and Remote Sensing – Focussing on high-performance computing aspects.
Appointed peer reviewer for multiple Journals and Press…
e.g. for the Oxford University Press „Briefings in Bioinformatics“ – Focussing on algorithms and high-performance computing aspects.
Certified in…
…HPC (High Performance Computing), Machine Learning (Stanford), PMP (Project Management Professional), SCRUM
Member of Societies…
…ACM & SPIE
GRANTS (thankfully received)
European Network on High Performance and Embedded Architecture and Compilation (HiPEAC), Berlin, Germany, Jan 2013 [Full Conference Fee Grant]
International Symposium on Code Generation and Optimization (CGO), San Francisco Bay Area, USA, Feb 2015 [ACM SIGPLAN Professional Activities Grant]
GPU Technology Conference (GTC), Silicon Valley (San José), USA, April 2016 [NVIDIA Full Conference Fee Grant]
Participant of Spring School on Mixed Integer Nonlinear Optimization and Applications, Tilburg University, March 10-13, 2015 [Full Travel Grant]
FOUNDER SCHOLARSHIP NRW
For innovative ideas. Monthly stipend for 1 year.
The NRW start-up grant gives you the opportunity to get your innovative business idea off the ground and to enter the start-up scene in your region. The Ministry of Economic Affairs, Innovation, Digitalization and Energy of the State of North Rhine-Westphalia supports every founder who is about to start up or is just starting out with a monthly stipend to facilitate the start in the world of entrepreneurs. In addition, you will have the opportunity to exchange ideas in start-up networks and be accompanied by individual coaching.
International Projects I worked on
RoKoRa
Secure human-robot collaboration using high-resolution radars
In the RoKoRa project, the aim is to exclude hazards of humans by robots. For this purpose, compact radar systems are used. They offer many advantages: radars operate independently of any lighting and are largely insensitive to environmental conditions. In addition, radar sensors can measure not only the distance to the sensor, but also the motion vector of the detected targets. The sensor system to be developed significantly improves personnel safety in human-robot collaboration, resulting in new degrees of freedom for a safe cooperation with larger robots and at higher distance-dependent execution speeds. In the project, SCAI develops a software component for environmental perception for situational analysis and decision-making using methods of machine learning.
Project duration: 1.7.2017 – 30.6.2020
WAVE
Simulation and inversion of wave fields
The aim of the project WAVE is to develop a toolbox that provides methods for the simulation and inversion of wave fields on high performance architectures. WAVE is funded by the BMBF (German National Ministry of Education and Research). Wave propagation plays an important role in many applications. In geophysical exploration, seismic and electromagnetic waves make geological layers and reservoirs visible. Seismology uses wave simulation to predict and estimate possible damages of earthquakes in high risk regions. Simulation models based on the theory of harmonic wave propagation are also used for the identification of tumors, for the localization of defects in materials, reducing the acoustic noise in a vehicle, or for the protection of places from electro-magnetic radiation.
Project duration:1.2.2016 – 31.6.2019
FAST
Find a Suitable Topology for exascale Applications
The project FAST funded by the BMBF will improve the initial job distribution of a scheduler and the re-scheduling on the basis of key performance indicators. From the perspective of a data center operator, it is uneconomical if not all components of a computer can be used by a running program and, in the worst case, still consume power. In the BMBF-funded project FAST, both the initial job distribution of a scheduler (system for distributing computing resources) and the re-scheduling based on key performance indicators (KPI) are to be improved. Rescheduling is the fine-tuning of schedules (a kind of time schedule) by migrating jobs between neighboring cluster nodes. The goal is to achieve a balanced distribution of the load and avoid any resource bottlenecks in the system by introducing scheduling phases.
Project duration: 1.1.2014 – 31.12.2016
MACH
The project MACH funded by the BMBF deals with the research for and development of concepts for the application of DSLs for the improvement of development and maintainability of computationally intensive algorithms on a multitude of hardware sytems.
Project duration: 1.11.2013 – 31.10.2016
ENHANCE
Enabling heterogeneous hardware acceleration using novel programming and scheduling models
The project aims for a better integration and simplified usage of heterogeneous computing resources within current and upcoming computing systems. Today, even personal computers may be used to solve medium-sized scientific computing problems in a cost-effective and energy-efficient way without accessing large and expensive super computing systems. SCAI develops methods for code analysis and automatic parallelization that enable the developer to exploit the maximum potential of a given hardware. The compiler plug-in PLUTO-SICA developed by SCAI includes, among other things, the automatic conversion of algorithmic code structures into those that work in parallel or vectorized mode and guarantee optimal memory access. The PLUTO-SICA development is funded by the BMBF in the ENHANCE project and is available as an optimization plug-in for the next release of the LLVM Compiler Suite.
Project duration: 1.11.2011 – 31.12.2013