mirrors/mesa - Frog Git

Commit Graph

Author	SHA1	Message	Date
Tim Rowley	95ed1c19bf	swr: allow alphatest without blend or logicop We need to compile a blend function when alphatest is enabled. Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2016-11-08 14:18:47 -06:00
Dave Airlie	bafc75b437	radv: emit correct last export when Z/stencil export is enabled I was getting a random GPU hang in the renderpass simple tests, it turns out sometimes radv emitted the wrong thing "last". This fixes the logic to emit Z/stencil last if they occur, and not mark a color output as last. Also this relies on the Z/STENCIL being the first two fragment outputs, which they are so yay. Fixes: dEQP-VK.renderpass.simple.color_depth (random hangs) Cc: "13.0" <mesa-stable@lists.freedesktop.org> Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-11-09 06:05:03 +10:00
Marek Olšák	bdd48e47c0	tgsi/scan: turn a huge if-else-if.. chain into a switch statement Reviewed-by: Brian Paul <brianp@vmware.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-08 17:56:42 +01:00
Marek Olšák	f864547fa9	tgsi/scan: fix images_buffers regression The first IF statement disabled the second one. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=98599 Reviewed-by: Brian Paul <brianp@vmware.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-08 17:56:42 +01:00
Jason Ekstrand	6b7cc8a9ec	anv: Document cmd_buffer_alloc_binding_table Some of the details of this function are very confusing and have a long history. We should document that history and this seems like the best place to do it. Signed-off-by: Jason Ekstrand <jason@jlekstrand.net>	2016-11-08 08:32:55 -08:00
Jason Ekstrand	406cd9d126	intel/blorp: Emit all the binding tables At least on Sky Lake, after emitting 3DSTATE_CONSTANT_*, you are required to re-emit the 3DSTATE_BINDING_TABLE_POINTERS packet for the corresponding stage. If you don't, double-buffering may fail and you may get the wrong constants. It turns out that you need to do this even if you have no push constants to speak of or else the next 3DSTATE_CONSTANT packet you emit for that stage may not work correctly. Signed-off-by: Jason Ekstrand <jason@jlekstrand.net> Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-11-08 08:32:55 -08:00
Jordan Justen	112a2ba276	i965/gen9: Allow sampling with hiz when supported For gen9+ this will indicate when we should allow hiz based sampling during rendering. Improves performance in : - Synmark's OglDeferred by 2.2% (n=20) - Synmark's OglShMapPcf by 0.44% (n=20) v2 by Ben: Add spec reference, and make it fix with some of the changes made on the previous patches Change the check from mt->aux_buf to mt->num_samples. The presence of an aux_buf isn't enough to determine there isn't a HiZ buffer to use. v3: It seems all depth surface end up with num_samples = 0 by default, so allow sampling from depth HiZ if num_samples <= 1. (Lionel) Allow sampling from HiZ only if all LOD are available from the HiZ buffer. (Lionel) Signed-off-by: Jordan Justen <jordan.l.justen@intel.com> (v1) Signed-off-by: Ben Widawsky <benjamin.widawsky@intel.com> (v2) Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> (v3) Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com>	2016-11-08 16:13:57 +00:00
Ben Widawsky	3b0c2bc417	i965/gen9: Add HiZ auxiliary buffer support The original functionality this patch introduces was authored by a patch from Ken (the commit subject was the same). Since I ended up changing so many patches in the code before this one, I had some non-trivial decisions to make, and I didn't feel it was appropriate to keeps Ken's name as author (mostly because he might not like what I've done). Ken's original patch was like 2 LOC :-) In either case, some credit needs to go to Ken, and to Jordan for a few small other changes in that original patch. v2: Back to a smaller diff now that ISL handles most of the actual programming (Lionel) Signed-off-by: Ben Widawsky <benjamin.widawsky@intel.com> (v1) Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> (v2) Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com>	2016-11-08 16:13:57 +00:00
Jordan Justen	c0f505c7ef	i965: Add function to indicate when sampling with hiz is supported Currently it indicates that this is never supported, but soon it will be supported for gen8+^w gen9+ v2 by Ben: - Explicitly disable aux_hiz for gen < 9 (with comment) - squashed in next patch to avoid unused and useless functions i965: Support sampling with hiz during rendering For gen8, we can sample from depth while using the hiz buffer. This allows us to sample depth without resolving from hiz to the depth texture. To do this we must resolve to hiz before drawing so we can use the hiz buffer to sample while rendering. Hopefully the hiz buffer will already be resolved in most cases because it was previously rendered, meaning the hiz resolve is a no-op. Note that this is still controlled by the intel_miptree_sample_with_hiz function, and we will enable hiz sampling for gen8 in a separate patch. Signed-off-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com> Signed-off-by: Jordan Justen <jordan.l.justen@intel.com> (v1) Signed-off-by: Ben Widawsky <benjamin.widawsky@intel.com> (v2) Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com>	2016-11-08 16:13:57 +00:00
Ben Widawsky	c53e9c9780	i965/miptree: Create a hiz mcs type This seems counter to the goal of consolidating hiz, mcs, and later ccs buffers. Unfortunately, hiz on gen6 is a thing the code supports, and this wart will be helpful to achieve that. Overall, I believe it does help unify AUX buffers on gen7+. I updated the size field which I introduced in the previous patch, even though we have no use for it. XXX: As I mentioned in the last patch, the height given to the MCS buffer allocation in intel_miptree_alloc_mcs() looks wrong, but I don't claim to fully understand how the MCS buffer is laid out. v2: rebase on master (Lionel) Signed-off-by: Ben Widawsky <benjamin.widawsky@intel.com> (v1) Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> (v2) Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com>	2016-11-08 16:13:57 +00:00
Ben Widawsky	36d1c555ed	i965: Drop the aux mt when not used This patch will preserve the BO & offset, and not the miptree for the aux_mcs buffer. Eventually it might make sense to pull put the sizing function in miptree creation, but for now this should be sufficient and not too hideous. v2: Save BO's offset too (Lionel) v3: Squash previous patch storing the size of the allocated aux buffer (Lionel) Fix memory leak with mcs_buf->bo (Lionel) Signed-off-by: Ben Widawsky <benjamin.widawsky@intel.com> (v1) Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> (v2) Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com>	2016-11-08 16:13:57 +00:00
Ben Widawsky	42db7ab179	i965/miptree: Directly gtt map the mcs buffer The next patch will change the map type, and this will make sure there are no regressions as a result of the other stuff. Since the miptree is newly created, I believe it is always safe to just map. It is possible to CPU map this buffer on LLC platforms (it additionally requires rounding up to tile size). I did experiment with that patch, and found no performance gains to be had. I've added in error handling while here. Generally GTT mapping is an operation which is highly unlikely to fail, but we may as well handle it when it does. v2: rebase on master (Lionel) v3: print out error if gtt mapping fails (Topi) Signed-off-by: Ben Widawsky <benjamin.widawsky@intel.com> (v1) Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> (v2) Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com>	2016-11-08 16:13:57 +00:00
Jordan Justen	0041169cac	i965: Wrap MCS miptree in intel_miptree_aux_buffer This will allow us to treat HiZ and MCS the same when using as an auxiliary surface buffer. v2: (Ben) Minor rebase conflict resolution. Rename mcs_buf to aux_buf to address upcoming change for hiz specific buffers. That second thing is essentially a squash of: i965/gen8: Use intel_miptree_aux_buffer for auxiliary buffer - which didn't need to be separate in my opinion. v3: rebase on master (Lionel) Signed-off-by: Jordan Justen <jordan.l.justen@intel.com> (v1) Signed-off-by: Ben Widawsky <benjamin.widawsky@intel.com>a (v2) Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> (v3) Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com>	2016-11-08 16:13:57 +00:00
Nicolai Hähnle	88f791db75	gallivm: fix [IU]MUL_HI regression This patch does two things: 1. It separates the host-CPU code generation from the generic code generation. This guards against accidently breaking things for radeonsi in the future. 2. It makes sure we actually use both arguments and don't just compute a square :-p Fixes a regression introduced by commit `29279f44b3` Cc: Roland Scheidegger <sroland@vmware.com> Reviewed-by: Roland Scheidegger <sroland@vmware.com>	2016-11-08 16:25:54 +01:00
Roland Scheidegger	3fa10ffb49	draw: use vectorized calculations for fetch Instead of doing all the math with scalars, use vectors. This means the overflow math needs to be done manually, albeit that's only really problematic for the stride/index mul, the rest has been pretty much moved outside the shader loop (albeit the mul could actually be optimized away too), where things are still scalar. Because llvm is complete fail with the zero-extend widening mul, roll our own even... To eliminate control flow in the main shader loop fetch, provide fake buffers (so index 0 is always valid to fetch). Still uses aos fetch though in the end - mostly because some more code would be needed to handle unaligned fetches in that path, and because for most formats it won't make a difference anyway (we generate some truly horrendous code for things like R16G16_something for instance). Instanced fetch however stays roughly the same as before, except that no longer the same element is fetched multiple times (I've seen a reduction of ~3 times in main shader loop size due to apparently llvm not being able to deduce it's really all the same with a couple instanced elements). Also, for elts gathering, use vectorized code as well - provide a fake elt buffer if there's no valid one bound. The generated shaders are smaller and faster to compile (not entirely sure about execution speed, but generally unless there's just single vertices to handle I would expect it to be faster - there's more opportunities for future improvements by using soa fetch). No piglit change. Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2016-11-08 03:41:26 +01:00
Roland Scheidegger	29279f44b3	gallivm: introduce 32x32->64bit lp_build_mul_32_lohi function This is used by shader umul_hi/imul_hi functions (and soon by draw). It's actually useful separating this out on its own, however the real reason for doing it is because we're using an optimized sse2 version, since the code llvm generates is atrocious (since there's no widening mul in llvm, and it does not recognize the widening mul pattern, so it generates code for real 64x64->64bit mul, which the cpu can't do natively, in contrast to 32x32->64bit mul which it could do). Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2016-11-08 03:41:26 +01:00
Anuj Phogat	b0554c25e7	i965: Add space before paren Signed-off-by: Anuj Phogat <anuj.phogat@gmail.com>	2016-11-07 16:13:57 -08:00
Anuj Phogat	501d608e56	i965: Remove unnecessary white space Signed-off-by: Anuj Phogat <anuj.phogat@gmail.com>	2016-11-07 16:13:57 -08:00
Anuj Phogat	329ae922bd	i965: Fix alpha-to-coverage and alpha test enabled checks Signed-off-by: Anuj Phogat <anuj.phogat@gmail.com> Reviewed-by: Ben Widawsky <ben@bwidawsk.net>	2016-11-07 16:13:02 -08:00
Anuj Phogat	a1bd2f6950	mesa: Add helper function _mesa_is_alpha_to_coverage_enabled() Signed-off-by: Anuj Phogat <anuj.phogat@gmail.com> Reviewed-by: Brian Paul <brianp@vmware.com> Reviewed-by: Ben Widawsky <ben@bwidawsk.net>	2016-11-07 16:13:02 -08:00
Anuj Phogat	0295c792b4	mesa: Add helper function _mesa_is_alpha_test_enabled() Signed-off-by: Anuj Phogat <anuj.phogat@gmail.com> Reviewed-by: Brian Paul <brianp@vmware.com> Reviewed-by: Ben Widawsky <ben@bwidawsk.net>	2016-11-07 16:13:02 -08:00
Anuj Phogat	7fed07766d	mesa: Use separate line for function return type Signed-off-by: Anuj Phogat <anuj.phogat@gmail.com> Reviewed-by: Brian Paul <brianp@vmware.com> Reviewed-by: Ben Widawsky <ben@bwidawsk.net>	2016-11-07 16:13:02 -08:00
Samuel Pitoiset	e32e5d214e	nvc0: simplify draw parameters upload for vertex shaders Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-11-07 22:50:17 +01:00
Steven Toth	381edca826	gallium/hud: protect against and initialization race In the event that multiple threads attempt to install a graph concurrently, protect the shared list. Signed-off-by: Steven Toth <stoth@kernellabs.com> Reviewed-by: Brian Paul <brianp@vmware.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-07 18:31:52 +01:00
Steven Toth	5a58323064	gallium/hud: close a previously opened handle We're missing the closedir() to the matching opendir(). Signed-off-by: Steven Toth <stoth@kernellabs.com> Reviewed-by: Brian Paul <brianp@vmware.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-07 18:31:52 +01:00
Steven Toth	6ffed08679	gallium/hud: fix a problem where objects are free'd while in use. Instead of trying to maintain a reference counted list of valid HUD objects, and freeing them accordingly, creating race conditions between unanticipated multiple threads, simply accept they're allocated once and never released until the process terminates. They're a shared resource between multiple threads, so accept they're always available for use. Signed-off-by: Steven Toth <stoth@kernellabs.com> Reviewed-by: Brian Paul <brianp@vmware.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-07 18:31:52 +01:00
Rob Clark	a5e733c6b5	mesa: drop current draw/read buffer when ctx is released This fixes a problem seen with gallium drivers vs android wallpaper. Basically, what happens is: EGLSurface tmpSurface = mEgl.eglCreatePbufferSurface(mEglDisplay, mEglConfig, attribs); mEgl.eglMakeCurrent(mEglDisplay, tmpSurface, tmpSurface, mEglContext); int[] maxSize = new int[1]; Rect frame = surfaceHolder.getSurfaceFrame(); glGetIntegerv(GL_MAX_TEXTURE_SIZE, maxSize, 0); mEgl.eglMakeCurrent(mEglDisplay, EGL_NO_SURFACE, EGL_NO_SURFACE, EGL_NO_CONTEXT); mEgl.eglDestroySurface(mEglDisplay, tmpSurface); ... check maxSize vs frame size and bail if needed ... mEglSurface = mEgl.eglCreateWindowSurface(mEglDisplay, mEglConfig, surfaceHolder, null); ... error checking ... mEgl.eglMakeCurrent(mEglDisplay, mEglSurface, mEglSurface, mEglContext); When the window-surface is created, it ends up with the same ptr address as the recently freed tmpSurface pbuffer surface. Which after many levels of indirection, results in st_framebuffer_validate() ending up with the same/old framebuffer object, and in the end never calling the DRIimageLoaderExtension::getBuffers(). Then in droid_swap_buffers(), the dri2_surf is still the old pbuffer surface (with dri2_surf->buffer being NULL, obviously, so when wallpaper app calls eglSwapBuffers() nothing gets enqueued to the compositor). Resulting in a black/blank background layer. Note that at the EGL layer, when the context is unbound, EGL drops it's references to the draw and read buffer as well. Signed-off-by: Rob Clark <robdclark@gmail.com> Tested-by: Robert Foss <robert.foss@collabora.com> Acked-by: Tapani Pälli <tapani.palli@intel.com>	2016-11-07 10:23:26 -05:00
Serge Martin	cc495055cd	clover: Add CL_PROGRAM_BINARY_TYPE support (CL1.2). v3 [Francisco Jerez]: Loosely based on Serge's v1 of this patch in order to avoid CL-specific enums in the clover module binary format. In addition to other changes made in v2: Represent the CL program binary type as the section type instead of adding a CL API-specific enum, check that the binary types of the input objects are valid during clLinkProgram(), pass section type as argument to build_module_library() instead of using separate function. Reviewed-by: Francisco Jerez <currojerez@riseup.net>	2016-11-06 15:56:54 +01:00
Serge Martin	05fcc73f08	clover: add missing clGetDeviceInfo CL1.2 queries Reviewed-by: Francisco Jerez <currojerez@riseup.net> Reviewed-by: Edward O'Callaghan <funfunctor@folklore1984.net> Reviewed-by: Vedran Miletić <vedran@miletic.net>	2016-11-06 15:56:49 +01:00
Samuel Pitoiset	8cc4a74971	nvc0: get rid of NVE4_COMPUTE_MP_PM_{A,B}_SIGSEL_XXX Instead, hardcode group sigsel because there are a bunch of unknown groups, especially on SM50/SM52. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com>	2016-11-05 19:28:25 +01:00
Samuel Pitoiset	a295364596	gm107/ir: emit RED instead of ATOM when no dst This is similar to NVC0 and GK110 emitters where we emit reduction operations instead of atomic operations when the destination is not used. Found after writing some tests which check if performance counters return the expected value. In that case, gred_count returned 0 on gm107 while at least gk106 returned the correct value. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-11-05 19:27:35 +01:00
Brian Paul	cfb5a9ab23	st/mesa: initialize members of glsl_to_tgsi_instruction in emit_asm() This fixes random crashes with MSVC release builds. It seems the members are implicitly initialized to zero with gcc, but not MSVC. In particular, the tex_offset_num_offset field was non-zero causing a loop over the NULL tex_offsets array to crash. Zero-init those fields and a few others to be safe. The regression began with `acc23b04cf` "ralloc: remove memset from ralloc_size". Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-05 12:09:40 -06:00
Mauro Rossi	0148313ea3	android: amd/common: add support for libmesa_amd_common Fixes the following building error introduced with commit `7115e56` and related amd/common dependencies: external/mesa/src/gallium/drivers/radeonsi/si_shader.c:6861: error: undefined reference to 'ac_is_sgpr_param' external/mesa/src/gallium/drivers/radeonsi/si_shader.c:6951: error: undefined reference to 'ac_is_sgpr_param' clang++: error: linker command failed with exit code 1 (use -v to see invocation) ninja: build stopped: subcommand failed. build/core/ninja.mk:148: recipe for target 'ninja_wrapper' failed make: *** [ninja_wrapper] Error 1 Signed-off-by: Marek Olšák <marek.olsak@amd.com>	2016-11-05 18:42:29 +01:00
Marek Olšák	0f72f7292a	winsys/radeon: don't call surface_best for FMASK Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=98518 Acked-by: Edward O'Callaghan <funfunctor@folklore1984.net>	2016-11-05 18:36:26 +01:00
Kenneth Graunke	0c17b0b6f0	mesa: Add linear ETC2/EAC to the compressed format list with ES3 compat. GL_ARB_ES3_compatibility brings ETC2/EAC formats to desktop GL. The meaning of the GL compressed format list is pretty vague - it's supposed to return formats for "general-purpose usage". (GL 4.2 deprecates the list because of this.) Basically everyone interprets this as "linear RGB/RGBA". ETC2/EAC meets that criteria, so while we shouldn't be required to add it to the list, there's also little harm in doing so, at least on platforms with native support. I doubt anyone is using this list for much anyway, so even on platforms without native support, it's probably not a big deal. Makes the following GL45-CTS.gtf43 tests pass: * GL3Tests.eac_compression_r11.gl_compressed_r11_eac * GL3Tests.eac_compression_rg11.gl_compressed_rg11_eac * GL3Tests.eac_compression_signed_r11.gl_compressed_signed_r11_eac * GL3Tests.eac_compression_signed_rg11.gl_compressed_signed_rg11_eac * GL3Tests.etc2_compression_rgb8.gl_compressed_rgb8_etc2 * GL3Tests.etc2_compression_rgb8_pt_alpha1.gl_compressed_rgb8_pt_alpha1_etc2 * GL3Tests.etc2_compression_rgba8.gl_compressed_rgba8_etc2 Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Eduardo Lima Mitev <elima@igalia.com>	2016-11-04 16:10:20 -07:00
Eric Anholt	283d4d18e5	vc4: Use Newton-Raphson on the 1/W write to fix glmark2 terrain. The 1/W was apparently not accurate enough, and we were getting sparklies in the distance. The closed driver also did a N-R step here. Cc: <mesa-stable@lists.freedesktop.org>	2016-11-04 15:34:38 -07:00
Eric Anholt	70fc3a941a	vc4: Make sure that vertex shader texture2D() calls use LOD 0. I noticed this while trying to debug glmark2 terrain (which does vertex shader texturing, but no mipmaps on its textures sampled from the VS).	2016-11-04 15:34:38 -07:00
Nicolai Hähnle	2c875158e2	radeonsi: fix vertex fetches for 2_10_10_10 formats The hardware always treats the alpha channel as unsigned, so add a shader workaround. This is rare enough that we'll just build a monolithic vertex shader. The SINT case cannot actually happen in OpenGL, but I've included it for completeness since it's just a mix of the other cases. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-04 21:30:18 +01:00
Nicolai Hähnle	322483f71b	st/mesa: fix the layer of VDPAU surface samplers A (latent) bug in VDPAU interop was exposed by commit `e5cc84dd43`. Before that commit, the st_vdpau code created samplers with first_layer == last_layer == 1 that the general texture handling code would immediately delete and re-create, because the layer does not match the information in the GL texture object. This was correct behavior at least in the DMABUF case, because the imported resource is supposed to have the correct offset already applied. In the non-DMABUF case, this was just plain wrong but apparently nobody noticed. After that commit, the state tracker assumes that an existing sampler is correct at all times. Existing samplers are supposed to be deleted when they may become invalid, and they will be created on-demand. This meant that the sampler with first_layer == last_layer == 1 stuck around, leading to rendering artefacts (on radeonsi), command stream failures (on r600), and assertions (in debug builds everywhere). This patch fixes the problem by simply not creating a sampler at all in st_vdpau_map_surface. We rely on the generic texture code to do the right thing, adding the layer_override to make the non-DMABUF case work. v2: add the layer_override Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=98512 Cc: 13.0 <mesa-stable@lists.freedesktop.org> Cc: Christian König <deathsimple@vodafone.de> Cc: Ilia Mirkin <imirkin@alum.mit.edu> Reviewed-by: Marek Olšák <marek.olsak@amd.com> (v1) Reviewed-by: Christian König <christian.koenig@amd.com>	2016-11-04 21:26:29 +01:00
Dave Airlie	d0d5f7600c	Revert "st/vdpau: use linear layout for output surfaces" This reverts commit `d180de3532`. This is a radeon specific hack that causes problems on nouveau when combined with the SHARED flag later. If radeonsi needs a fix for this, please fix it in the driver. [chk] Using linear surfaces for this makes sense because tilling isn't beneficial and the surfaces can potentially be shared with other GPUs using the VDPAU OpenGL interop. [airlied] I think we need a flag that isn't SHARED/LINEAR that is more SHARED_OTHER_GPU. [mareko] Does radeonsi need PIPE_BIND_VIDEO_DECODE_OUTPUT that it would translate into linear ? [mareko] My only concern is decoding performance. If the decoder works in 64x1 blocks, tiling will hurt. That's the theory. I don't know how the decoder works. Cc: 12.0 13.0 <mesa-stable@lists.freedesktop.org> Acked-by: Christian König <christian.koenig@amd.com> Signed-off-by: Dave Airlie <airlied@redhat.com> Tested-by: Ilia Mirkin <imirkin@alum.mit.edu> Tested-by: Nayan Deshmukh <nayan26deshmukh@gmail.com> (I+A)	2016-11-04 15:04:21 +00:00
Marek Olšák	00baaa4752	radeonsi: fix an assertion failure in si_decompress_sampler_color_textures This fixes a crash in Deus Ex: Mankind Divided. Release builds were unaffected, so it's not too serious. Cc: 11.2 12.0 13.0 <mesa-stable@lists.freedesktop.org> Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-04 11:30:47 +01:00
Marek Olšák	64c2593a5c	glx: make interop ABI visible again This was broken when the GLAPI use was removed from mesa_glinterop.h. Cc: 12.0 13.0 <mesa-stable@lists.freedesktop.org> Acked-by: Alex Deucher <alexander.deucher@amd.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> Reviewed-by: Emil Velikov <emil.velikov@collabora.com>	2016-11-04 11:30:47 +01:00
Marek Olšák	ee39d4456e	egl: make interop ABI visible again This was broken when the GLAPI use was removed from mesa_glinterop.h. Cc: 12.0 13.0 <mesa-stable@lists.freedesktop.org> Acked-by: Alex Deucher <alexander.deucher@amd.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> Reviewed-by: Emil Velikov <emil.velikov@collabora.com>	2016-11-04 11:30:47 +01:00
Marek Olšák	bf51b45313	egl: use util/macros.h I need the definition of PUBLIC. Cc: 12.0 13.0 <mesa-stable@lists.freedesktop.org> Acked-by: Alex Deucher <alexander.deucher@amd.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> Reviewed-by: Emil Velikov <emil.velikov@collabora.com>	2016-11-04 11:30:47 +01:00
Nicolai Hähnle	84a74be9e4	radeonsi: enable GLSL 4.50 Reviewed-by: Edward O'Callaghan <funfunctor@folklore1984.net> Reviewed-by: Dave Airlie <airlied@redhat.com>	2016-11-04 10:33:50 +01:00
Nicolai Hähnle	e4b378800e	st/glsl_to_tgsi: fix dvec[34] loads from SSBO When splitting up loads, we have to add 16 bytes to the offset for the high components, just like already happens for stores. Fixes arb_gpu_shader_fp64@shader_storage@layout-std140-fp64-shader. Cc: 13.0 <mesa-stable@lists.freedesktop.org> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-04 10:31:02 +01:00
Nicolai Hähnle	aef7eb4cac	glsl/cache: correct asprintf error handling From the manpage of asprintf: "If memory allocation wasn't possible, or some other error occurs, these functions will return -1, and the contents of strp are undefined." Reviewed-by: Iago Toral Quiroga <itoral@igalia.com> Reviewed-by: Edward O'Callaghan <funfunctor@folklore1984.net>	2016-11-04 10:28:08 +01:00
Michel Dänzer	8ce7ef75f5	gallium/radeon: Multiply bpe by nsamples in surf_winsys_to_drm For symmetry with surf_drm_to_winsys. Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-04 16:51:18 +09:00
Michel Dänzer	356458363d	gallium/radeon: Use flags parameter in radeon_winsys_surface_init Fixes valgrind warnings about surf_ws->flags being uninitialized while starting X. Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-04 16:49:39 +09:00
Michel Dänzer	6f844a30c1	gallium/radeon: Only convert stencil info if RADEON_SURF_SBUFFER is set Fixes valgrind warnings about using uninitialized memory when starting X. Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-04 16:48:59 +09:00
Michel Dänzer	38fb9aa1aa	gallium/radeon: Only loop up to last_level for drm<->winsys conversion Fixes spurious assertion failure in surf_level_drm_to_winsys when starting X, due to processing a miplevel which was never initialized. Fixes: `e9c76eeeaa` ("gallium/radeon: remove radeon_surf_level::pitch_bytes") Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-04 16:47:43 +09:00
Tapani Pälli	1e3f7bfc9a	anv: use limits.h instead of deprecated/obsolete values.h Mesa uses limits.h elsewhere, and this makes is possible to compile anv_allocator.c on Android. Signed-off-by: Tapani Pälli <tapani.palli@intel.com> Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com>	2016-11-04 08:35:43 +02:00
Eric Anholt	80157466cd	vc4: Add miptree/texture state support for ETC1 compressed textures. The format isn't flagged as enabled at runtime yet, because we need kernel validation support.	2016-11-03 18:42:58 -07:00
Eric Anholt	bedb996087	vc4: Fix use of undefined values since the ralloc zeroing changes. reralloc() no longer zeroes the new contents, so switch to using rzalloc_array() instead.	2016-11-03 18:42:58 -07:00
Eric Anholt	49936364e4	nir: Make sure to set the texsrc type in nir drawpixels/bitmap lowering. We were leaving an undefined value since the ralloc zeroing changes. Fixes nir_validate() failures on vc4. v2: Fix the color-index case of drawpixels as well. Reviewed-by: Rob Clark <robdclark@gmail.com> (v1)	2016-11-03 18:42:58 -07:00
Roland Scheidegger	572a952126	draw: fix undefined input handling some more... Previous fixes were incomplete - some code still iterated through the number of elements provided by velem layout instead of the number stored in the key (which is the same as the number defined by the vs). And also actually accessed the elements from the layout directly instead of those in the key. This mismatch could still cause crashes. (Besides, it is a very good idea to only use data stored in the key anyway.) v2: move null format check, remove now unnecessary function parameter, some minor prettify Reviewed-by: Jose Fonseca <jfonseca@vmware.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-04 01:48:22 +01:00
Brian Paul	f4dd3bde37	gallium/hud: call fflush() after printing error messages For Windows. Otherwise, we don't see the message until the program exits. Reviewed-by: Charmaine Lee <charmainel@vmware.com>	2016-11-03 14:29:23 -06:00
Brian Paul	260d951486	svga: move svga_mark_surfaces_dirty() prototype to svga_surface.h Trivial.	2016-11-03 14:29:23 -06:00
Brian Paul	c96f63cac2	svga: whitespace / formatting clean-up in svga_context.c Trivial.	2016-11-03 14:29:23 -06:00
Brian Paul	1691e29e62	svga: collect stats for time spent in svga_context_finish() This should have appeared with commit "svga: add guest statistic gathering interface" from August 4, but was somehow lost.	2016-11-03 14:29:23 -06:00
Charmaine Lee	8a195e2fd5	svga: invalidate new surface before it is bound to a render target view Invalidate a "new" surface before it is bound to a render target view or depth stencil view in order to avoid the unnecessary host side copy of the surface data before it is rendered to. Note that, recycled surface is already invalidated before it is reused. Reviewed-by: Brian Paul <brianp@vmware.com>	2016-11-03 14:29:23 -06:00
Charmaine Lee	06bba2452f	Revert "svga: use untyped surface formats in most cases" Using untyped surface formats causes huge performance degradation on Fusion. This reverts commit `eb0ced74f6` until the backend has a better solution to address typeless surface formats.	2016-11-03 14:29:23 -06:00
Charmaine Lee	f2eec4e829	svga: allow quad blit for more formats Currently blitter will fail if the blit format is different and view-incompatible to the resource format. Instead of punting to software blit which will stall the pipeline, we will create temporary resource to allow blitter to work. Fixes piglit test arb_copy_image-formats. Also tested with MTT piglit, glretrace. Reviewed-by: Brian Paul <brianp@vmware.com>	2016-11-03 14:29:22 -06:00
Charmaine Lee	4bd5ce853b	svga: create BGRX render target view for BGRX_UNORM surface Currently we adjust the view format when we are asked to create a BGRA render target view for BGRX surface. But we only look for SVGA3D_B8G8R8X8_TYPELESS surface format. With this patch, we will also check for SVGA3D_B8G8R8X8_UNORM surface format, and use SVGA3D_B8G8R8X8_UNORM as the view format for that case. Reviewed-by: Brian Paul <brianp@vmware.com>	2016-11-03 14:29:22 -06:00
Charmaine Lee	0d221fcd40	svga: add a helper function to check for typeless format This patch adds a helper function svga_format_is_typeless() which returns TRUE if the specified format is typeless. Reviewed-by: Brian Paul <brianp@vmware.com>	2016-11-03 14:29:22 -06:00
Brian Paul	d451421bca	svga: add SVGA_NEW_FRAME_BUFFER to svga_hw_tss_binding state atom We may need to re-emit texture bindings when the framebuffer state changes. In particular, emitting the texture binding can also involve updating a texture from its backing copy during sampler view validation. The backing copy is made during framebuffer validation. This helps to fix an issue with Photoshop on VGPU9 (VMware bug 1723971). Reviewed-by: Charmaine Lee <charmainel@vmware.com>	2016-11-03 14:29:22 -06:00
Charmaine Lee	ec138d6237	svga: allow copy_region if sample counts match With this patch, we will allow blit with copy_region if the source and destination textures have the same sample counts. Fixes failures with piglit tests spec@arb_texture_float@multisample-formats 2 gl_arb_texture_float spec@arb_texture_rg@multisample-formats 2 gl_arb_texture_rg-float Reviewed-by: Brian Paul <brianp@vmware.com>	2016-11-03 14:29:22 -06:00
Charmaine Lee	a2d49c4b46	svga: set rendered-to flag after updating the texture using PredCopyRegion This patch sets the rendered-to flag for the subresource after it is updated using the PredCopyRegion command. This is to ensure that the GB surface will be sync up properly before it will be directly mapped to. Tested with MTT piglit, glretrace. Reviewed-by: Brian Paul <brianp@vmware.com>	2016-11-03 14:29:22 -06:00
Charmaine Lee	59f14563a3	svga: add can_use_upload flag This patch adds a flag "can_use_upload" to svga_texture structure to avoid some checking of the upload availability at each transfer map time. Tested with Lightsmark2008, Tropics, MTT glretrace, piglit. Reviewed-by: Brian Paul <brianp@vmware.com>	2016-11-03 14:29:22 -06:00
Charmaine Lee	3dfb4243bd	svga: fix texture upload path condition As Thomas suggested, we'll first try to map directly to a GB surface. If it is blocked, then we'll use texture upload buffer. Also if a texture is already "rendered to", that is, the GB surface is already out of sync, then we'll use the texture upload buffer to avoid syncing the GB surface. Tested with Lightsmark2008, Tropics, MTT piglit, glretrace. Reviewed-by: Brian Paul <brianp@vmware.com>	2016-11-03 14:29:22 -06:00
Charmaine Lee	4750c4e543	svga: set rendered_to flag with texture uploaded using TransferFromBuffer command This patch sets the rendered_to flag for the texture subresource that is uploaded using the TransferFromBuffer command. This is to ensure that the subresource will be read back or invalidated before it will be directly mapped to. This makes sure that the content of the GB surface will not be accidentally overwritten by the device at suspend/resume time. Reviewed-by: Brian Paul <brianp@vmware.com>	2016-11-03 14:29:22 -06:00
Neha Bhende	03e1b7cacd	svga: Add render_condition boolean flag in struct svga_context set render_condition flag when driver performs conditional rendering. Blit using DXPredCopyRegion command gets affected by conditional rendering so We should check this flag while performing blit operation Tested with piglit tests. v2: As per Charmaine's comment, setting render_condition flag if svga_query is valid. Tested with pigit tests. Reviewed-by: Brian Paul <brianp@vmware.com> Reviewed-by: Charmaine Lee <charmainel@vmware.com>	2016-11-03 14:29:22 -06:00
Neha Bhende	2cff6f4512	svga: Allow DXPredCopyRegion for depth_and_stencil formats. DXPredCopyRegion supports copy between src and dst for depth_and_stencil formats if src and dst have same formats. tested ith piglit v2: As per Brian's comment, allow DXPredCopyRegion for depth+stencil buffers if the blit mask is PIPE_MASK_ZS. Tested with piglit tests and added new piglit test arb_framebuffer_object-depth-stencil-blit to test this particular testcase. Reviewed-by: Brian Paul <brianp@vmware.com>	2016-11-03 14:29:22 -06:00
Neha Bhende	9a9627a791	svga: fix memory leak in svga_clear_texture() Piglit tests which uses arb_clear_texture extension, have memory leak issue. pipe_surface created in svga_clear_texture() was not deleted which happens to be the cause for memory leak. tested all arb_clear_texture-* piglit tests with valgrid. Reviewed-by: Brian Paul <brianp@vmware.com> Reviewed-by: Charmaine Lee <charmainel@vmware.com>	2016-11-03 14:29:22 -06:00
Thomas Hellstrom	d787ce7288	svga: Implement the pipe clear_render_target functionality v2 v2: Accounted for the fact that svga_try_clear_render_target also honors conditional rendering. Testing done: Excercised all functions in a separate feature branch. Forced emission of conditional rendering commands when necessary. Signed-off-by: Thomas Hellstrom <thellstrom@vmware.com> Reviewed-by: Brian Paul <brianp@vmware.com>	2016-11-03 14:29:22 -06:00
Charmaine Lee	76f5f76468	svga: add SVGA_3D_CMD_INVALIDATE_GB_SURFACE support This command will be used in a subsequent patch to invalidate a surface. Reviewed-by: Brian Paul <brianp@vmware.com>	2016-11-03 14:29:22 -06:00
Francisco Jerez	f3d387867f	nir: Flip gl_SamplePosition in nir_lower_wpos_ytransform(). Assuming the hardware is set up to use a screen coordinate system flipped vertically with respect to the GL's window coordinate system, the SYSTEM_VALUE_SAMPLE_POS vector will also be flipped vertically with respect to the value expected by the GL, so we need to give it the same treatment as gl_FragCoord. Fixes the following CTS tests on i965: ES31-CTS.functional.shaders.multisample_interpolation.interpolate_at_offset.at_sample_position.default_framebuffer ES31-CTS.functional.shaders.sample_variables.sample_pos.correctness.default_framebuffer when run with any multisample configuration, e.g. rgba8888d24s8ms4. Cc: <mesa-stable@lists.freedesktop.org> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2016-11-03 11:46:44 -07:00
Nanley Chery	faab6a0f18	isl: Only allow Y-tiling for ASTC textures Signed-off-by: Nanley Chery <nanley.g.chery@intel.com> Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-11-03 11:22:58 -07:00
Nanley Chery	1625d911d7	anv/blorp: Don't create linear ASTC surfaces for buffers Such a surface is not possible on our hardware. Without this change, ISL surface creation would fail with the next patch. Signed-off-by: Nanley Chery <nanley.g.chery@intel.com> Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-11-03 11:22:58 -07:00
Nanley Chery	bb550e2977	anv/formats: Disallow linear ASTC textures Signed-off-by: Nanley Chery <nanley.g.chery@intel.com> Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-11-03 11:22:58 -07:00
Nanley Chery	80de528c7e	anv/formats: Disallow 1D compressed textures Signed-off-by: Nanley Chery <nanley.g.chery@intel.com> Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-11-03 11:22:58 -07:00
Chris Wilson	b4001af174	i965: Use rzalloc for cfg_t Valgrind reports that we use cfg.cycle_count uninitialised, so zero the cfg_t on construction. Fixes: `52d2b28f7f` ("ralloc: use rzalloc where it's necessary") Signed-off-by: Chris Wilson <chris@chris-wilson.co.uk> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 11:16:05 -07:00
Nicolai Hähnle	27bd9c0f0a	pipe-loader: add libamd_common for radeonsi This fixes a build regression of commit `7115e56c21`. Sorry for the breakage, this second location for link dependencies escaped my build tests. Bugzilla: https://patchwork.freedesktop.org/patch/119816/ Tested-by: Dieter Nützel <Dieter@nuetzel-hh.de>	2016-11-03 16:54:55 +01:00
Andreas Boll	f792f0687f	glx/windows: Add wgl.h to the sources list Otherwise it won't be picked in the tarball and the build will fail. Fixes: `533b3530c1` ("direct-to-native-GL for GLX clients on Cygwin ("Windows-DRI")") Cc: "13.0" <mesa-stable@lists.freedesktop.org> Signed-off-by: Andreas Boll <andreas.boll.dev@gmail.com> Reviewed-by: Jon Turney <jon.turney@dronecode.org.uk>	2016-11-03 11:38:04 +01:00
Tapani Pälli	979ec2cf75	i965: use rzalloc instead of calloc in brwNewProgram commit `cc6aa1d161` changed to using rzalloc for gl_program creation but one instance for program creation was still using calloc. Signed-off-by: Tapani Pälli <tapani.palli@intel.com> Reviewed-by: Juan A. Suarez <jasuarez@igalia.com> Reviewed-by: Iago Toral Quiroga <itoral@igalia.com> Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com>	2016-11-03 12:09:40 +02:00
Nicolai Hähnle	908f92ad1f	radeonsi: generate GS prolog to (partially) fix triangle strip adjacency rotation Fixes GL45-CTS.geometry_shader.adjacency.adjacency_indiced_triangle_strip and others. This leaves the case of triangle strips with adjacency and primitive restarts open. It seems that the only thing that cares about that is a piglit test. Fixing this efficiently would be really involved, and I don't want to use the hammer of degrading to software handling of indices because there may well be software that uses this draw mode (without caring about the precise rotation of triangles). v2: - skip the GS prolog entirely if workaround is not needed - only check for TES (TES is always non-null when tessellation is used) Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:11:24 +01:00
Nicolai Hähnle	ffe4e829b0	radeonsi: remove si_shader_context::is_gs_copy_shader It has become redundant. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:53 +01:00
Nicolai Hähnle	3b2516721b	radeonsi: make the GS copy shader owned by the GS selector The copy shader only depends on the selector. This change avoids creating separate code paths for monolithic vs. non-monolithic geometry shaders. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:50 +01:00
Nicolai Hähnle	9c6f7d66dc	radeonsi: si_shader_vs only depends on the GS selector Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:48 +01:00
Nicolai Hähnle	693435d846	radeonsi: si_vgt_gs_mode only depends on the selector Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:45 +01:00
Nicolai Hähnle	2e1fb7e7fc	radeonsi: make si_generate_gs_copy_shader usable as a standalone function It really only depends on the shader selector. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:42 +01:00
Nicolai Hähnle	ba5de0d034	radeonsi: unify the si_compile_* functions for prologs and epilogs Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:37 +01:00
Nicolai Hähnle	aa9583b0da	radeonsi: get rid of no_{prolog,epilog} Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:34 +01:00
Nicolai Hähnle	75503b1904	radeonsi: get rid of si_llvm_emit_fs_epilogue It is no longer used. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:31 +01:00
Nicolai Hähnle	611510038a	radeonsi: get rid of get_interp_param Replace by a simple LLVMGetParam, since ctx->no_prolog is always false. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:29 +01:00
Nicolai Hähnle	3f4439b6ba	radeonsi: get rid of select_interp_param The condition !ctx->no_prolog is now always true. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:26 +01:00
Nicolai Hähnle	858ac2228f	radeonsi: use TCS epilog for monolithic shaders For fixed function TCS, we keep the copying of VS outputs to TES inputs inside the main function; the call to si_copy_tcs_inputs is moved accordingly. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:23 +01:00
Nicolai Hähnle	3f1be54e53	radeonsi: extract si_build_tcs_epilog_function Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:20 +01:00
Nicolai Hähnle	be6e31c6a0	radeonsi: use VS epilog for monolithic TES Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:17 +01:00
Nicolai Hähnle	06dcb2d2a9	radeonsi: use VS prolog and epilog for monolithic shaders Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:14 +01:00
Nicolai Hähnle	f9daa2f470	radeonsi: extract si_build_vs_{prolog,epilog}_function Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:12 +01:00
Nicolai Hähnle	6f37e992a3	radeonsi: use PS prolog for monolithic shaders Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:09 +01:00
Nicolai Hähnle	15dd332e6a	radeonsi: set num_input_vgprs for fragment shaders in create_function So that the prolog generated for monolithic fragment shaders will have the right signature. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:05 +01:00
Nicolai Hähnle	fec7ced211	radeonsi: extract si_build_ps_prolog_function Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:02 +01:00
Nicolai Hähnle	7115e56c21	radeonsi: use PS epilog for monolithic shaders Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:07:00 +01:00
Nicolai Hähnle	bf86c56594	radeonsi: extract si_build_ps_epilog_function Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:06:57 +01:00
Nicolai Hähnle	0b9bba7f6c	radeonsi: pass the function name to si_llvm_create_func We will use multiple functions in one module, so they should have different names. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:06:54 +01:00
Nicolai Hähnle	96d60dd9ee	radeonsi: split is_monolithic into no_prolog and no_epilog This helps to achieve a gradual transition towards building monolithic shaders via inlining. no_prolog and no_epilog will be removed by the end of the series, separate_prolog remains in use to control the PS input mapping. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:06:50 +01:00
Nicolai Hähnle	8db9d915cd	radeonsi: free data structures when shader compiles fail Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:06:47 +01:00
Nicolai Hähnle	4c1504af6a	radeonsi: move main TGSI translation into its own function The idea is that adding prolog and epilog code will be pulled out into the caller. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:06:44 +01:00
Nicolai Hähnle	23dfb688ba	radeonsi: add always-inline pass to si_llvm_finalize_module Change the pass manager as well, since this is a module-level pass. No noticeable run-time difference on shader-db. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:06:42 +01:00
Nicolai Hähnle	4ada1dabc4	radeonsi: fix signature of export intrinsic in VS epilog The incompatible signature becomes an issue when the VS epilog gets merged with the main vertex shader at the IR level. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:06:33 +01:00
Nicolai Hähnle	899b2f24a4	radeonsi: link against amd_common Reviewed-by: Dave Airlie <airlied@redhat.com> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:06:30 +01:00
Nicolai Hähnle	908100cfae	amd/common: add ac_is_sgpr_param helper Reviewed-by: Dave Airlie <airlied@redhat.com> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:06:27 +01:00
Nicolai Hähnle	2ff5df8f50	amd/common: build also for gallium drivers At least when LLVM is used, which is basically always (unless you're only building r600 without OpenCL). Reviewed-by: Dave Airlie <airlied@redhat.com> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:06:24 +01:00
Nicolai Hähnle	8eabee9ec0	amd/common: move llvm helper prototype to ac_llvm_util.h Reviewed-by: Dave Airlie <airlied@redhat.com> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-03 10:05:46 +01:00
Nicolai Hähnle	37d646c1b3	glsl: fix lowering of UBO references of named blocks When a UBO reference has the form block_name.foo where block_name refers to a block where the first member has a non-zero offset, the base offset was incorrectly added to the reference. Fixes an assertion triggered in debug builds by GL45-CTS.enhanced_layouts.uniform_block_layout_qualifier_conflict. That test doesn't properly check for correct execution in this case, so I am also going to send out a piglit test. Cc: 13.0 <mesa-stable@lists.freedesktop.org> Reviewed-by: Iago Toral Quiroga <itoral@igalia.com>	2016-11-03 09:57:25 +01:00
Kenneth Graunke	8df4aebc94	glsl: Update deref types when resizing implicitly sized arrays. At link time, we resolve the size of implicitly sized arrays. When doing so, we update the type of the ir_variables. However, we neglected to update the type of ir_dereference nodes which reference those variables. It turns out array_resize_visitor (for GS/TCS/TES interface array handling) already did 2/3 of the cases for this, so we can simply refactor the code and reuse it. This fixes: GL45-CTS.shader_storage_buffer_object.basic-syntax GL45-CTS.shader_storage_buffer_object.basic-syntaxSSO which have an SSBO containing an implicitly sized array, followed by some other members. setup_buffer_access uses the dereference types to compute offsets to fields, and it had a stale type where the implicitly sized array's length was still 0 instead of the actual length. While we're here, we can also fix update_array_sizes to properly update deref types as well, fixing a FINISHME from 2010. Cc: mesa-stable@lists.freedesktop.org Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com>	2016-11-03 01:42:37 -07:00
Timothy Arceri	d2861d682a	mesa/glsl: delete previously linked shaders earlier when linking This moves the delete linked shaders call to _mesa_clear_shader_program_data() which makes sure we delete them before returning due to any validation problems. It also reduces some code duplication. From the OpenGL 4.5 Core spec: "If LinkProgram failed, any information about a previous link of that program object is lost. Thus, a failed link does not restore the old state of program. ... If one of these commands is called with a program for which LinkProgram failed, no error is generated unless otherwise noted. Implementations may return information on variables and interface blocks that would have been active had the program been linked successfully. In cases where the link failed because the program required too many resources, these commands may help applications determine why limits were exceeded." Therefore it's expected that we shouldn't be able to query the program that failed to link and retrieve information about a previously successful link. Before this change the linker was doing validation before freeing the previously linked shaders and therefore could exit on failure before they were freed. This change also fixes an issue in compat profile where a program with no shaders attached is expect to fall back to fixed function but was instead trying to relink IR from a previous link. Reviewed-by: Tapani Pälli <tapani.palli@intel.com> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=97715 Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-11-03 11:58:53 +11:00
Timothy Arceri	903e5eae97	nir: fix nir_shader_clone() and nir_sweep() These were broken in `e1af20f18a` when the info field in nir_shader was turned into a pointer. Clone was copying the pointer rather than the data and nir_sweep was cleaning up shader_info rather than claiming it. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-11-03 10:39:13 +11:00
Timothy Arceri	f304aca542	mesa: move shader_info to the start of gl_program This will allow use to use ralloc_parent() on the info field and fix a regression in nir_sweep() caused by `e1af20f18a`. This is intended to be a temporary requirement that will be removed when we finish separating shader_info from nir_shader. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-11-03 10:39:13 +11:00
Timothy Arceri	cc6aa1d161	st/mesa/r200/i915/i965: use rzalloc() to create gl_program This allows us to use ralloc_parent() to see which data structure owns shader_info which allows us to fix a regression in nir_sweep(). This will also allow us to move some fields from gl_linked_shader to gl_program, which will allow us to do some clean-ups like storing gl_program directly in the CurrentProgram array in gl_pipeline_object enabling some small validation optimisations at draw time. Also it is error prone to depend on the gl_linked_shader for programs in current use because a failed linking attempt will free infomation about the current program. In i965 we could be trying to recompile a shader variant but may have lost some required fields. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-11-03 10:39:13 +11:00
Samuel Pitoiset	548b5fee6b	nv50,nvc0: stop limiting the number of active queries to 1 This limitation was initially here because AMD_performance_monitor doesn't allow to expose the real number of hardware counters. But this actually really annoying when profiling with qapitrace. Anyways, performance counters are mostly for developers and failures are expected if you try to monitor more queries than supported. This breaks amd_performance_monitor_measure but it's expected. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com>	2016-11-02 23:42:09 +01:00
Samuel Pitoiset	b6137f226c	nvc0: add new warp_nonpred_execution_efficiency metric on SM35 Event not_predicated_off_thread_inst_executed is SM35+. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-11-02 23:35:49 +01:00
Samuel Pitoiset	98a382d013	nvc0: add missing metric-issue_slot on SM35 Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-11-02 23:35:46 +01:00
Samuel Pitoiset	c32d7175aa	nvc0: do not expose metric-inst_issued twice on SM35 Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-11-02 23:35:44 +01:00
Samuel Pitoiset	524703da58	nvc0: add new warp_execution_efficiency metric on SM30+ Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-11-02 23:35:42 +01:00
Samuel Pitoiset	51fe48660a	nvc0: respect 80-chars for perf metrics descriptions Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-11-02 23:35:39 +01:00
Samuel Pitoiset	b58d85bac8	nvc0: sort performance metrics alphabetically Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-11-02 23:35:28 +01:00
Fredrik Höglund	e7b9c5eb74	radv: add support for anisotropic filtering on VI+ Ported from radeonsi. Cc: "13.0" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-11-03 08:27:21 +10:00
Dave Airlie	73592b9284	radv: fix dual source blending Dolphin tried to use this, but we hadn't had any tests for it properly. All that is required is the shader output format needs to be set for 0 and 1 exports. Cc: "13.0" <mesa-stable@lists.freedesktop.org> Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-11-03 08:26:51 +10:00
Samuel Pitoiset	1d75d681d3	nv50: add missing draw_calls_indexed driver stat Spotted when glancing at the VBO push code. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-11-02 21:11:57 +01:00
Adam Jackson	afaaf623d4	glx/glvnd: Use bsearch() in FindGLXFunction instead of open-coding it Reviewed-by: Eric Engestrom <eric.engestrom@imgtec.com> Signed-off-by: Adam Jackson <ajax@redhat.com>	2016-11-02 14:52:43 -04:00
Adam Jackson	8bca8d89ef	glx/glvnd: Fix dispatch function names and indices As this array was not actually sorted, FindGLXFunction's binary search would only sometimes work. Cc: "13.0" <mesa-stable@lists.freedesktop.org> Reviewed-by: Eric Engestrom <eric.engestrom@imgtec.com> Signed-off-by: Adam Jackson <ajax@redhat.com>	2016-11-02 14:52:38 -04:00
Adam Jackson	deb0eb1660	glx/glvnd: Don't modify the dummy slot in the dispatch table Cc: "13.0" <mesa-stable@lists.freedesktop.org> Reviewed-by: Eric Engestrom <eric.engestrom@imgtec.com> Signed-off-by: Adam Jackson <ajax@redhat.com>	2016-11-02 14:52:31 -04:00
Jason Ekstrand	71cc1e188d	anv/pipeline: Properly cache prog_data::param Before we were caching the prog data but we weren't doing anything with brw_stage_prog_data::param so anything with push constants wasn't getting cached properly. This commit fixes that. Signed-off-by: Jason Ekstrand <jason@jlekstrand.net> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=98012 Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-11-02 09:32:28 -07:00
Jason Ekstrand	ff3185e3ba	anv/pipeline: Put actual pointers in anv_shader_bin While we can simply calculate offsets to get to things such as the prog_data and the key, it's much more user-friendly if there are just pointers. Also, it's a bit more fool-proof. While we're at it, we rework the pipeline cache API to use the brw_stage_prog_data type directly. Signed-off-by: Jason Ekstrand <jason@jlekstrand.net> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=98012 Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com> Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-11-02 09:32:22 -07:00
Jason Ekstrand	4306c10a88	intel/blorp: Pass a brw_stage_prog_data to upload_shader Signed-off-by: Jason Ekstrand <jason@jlekstrand.net> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=98012 Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com> Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-11-02 09:32:19 -07:00
Jason Ekstrand	058304f081	intel/blorp: Use wm_prog_data instead of hand-rolling our own Signed-off-by: Jason Ekstrand <jason@jlekstrand.net> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=98012 Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com> Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-11-02 09:32:15 -07:00
Jason Ekstrand	a5f8ff6ca1	anv: Better handle return codes from anv_physical_device_init The case where we just want the loop to continue is INCOMPATIBLE_DRIVER because that simply means that whatever FD we opened isn't a supported Intel chip. Other error codes such as OUT_OF_HOST_MEMORY are actual errors and we should be returning early in that case. Signed-off-by: Jason Ekstrand <jason@jlekstrand.net> Reviewed-by: Dave Airlie <airlied@redhat.com> Reviewed-by: Eric Engestrom <eric.engestrom@imgtec.com> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-11-02 09:26:41 -07:00
Jason Ekstrand	daeb21e478	vulkan/wsi/x11: Clean up connections in finish_wsi Signed-off-by: Jason Ekstrand <jason@jlekstrand.net> Reviewed-by: Dave Airlie <airlied@redhat.com> Reviewed-by: Eric Engestrom <eric.engestrom@imgtec.com> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-11-02 09:26:36 -07:00
Jason Ekstrand	fc0e9e3e40	vulkan/wsi/x11: Better handle wsi_x11_connection_create failure Without this fix, the function would still end up returning NULL but it would put that NULL connection in the hash table which would be bad. Signed-off-by: Jason Ekstrand <jason@jlekstrand.net> Reviewed-by: Dave Airlie <airlied@redhat.com> Reviewed-by: Eric Engestrom <eric.engestrom@imgtec.com> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-11-02 09:25:57 -07:00
Nicolai Hähnle	1ef505bb02	glsl: compute lvalues of [in]out parameters before inlined function body This is required when an out argument involves an array index that is either a global variable modified by the function or another out argument in the same function call. Fixes the shaders/out-parameter-indexing/vs-inout-index-inout-* tests. v2: - modify the ir_dereference_array nodes in place - use ir_hierarchical_visitor v3: use base_ir (Ian Romanick) Reviewed-by: Ian Romanick <ian.d.romanick@intel.com>	2016-11-02 12:32:47 +01:00
Nicolai Hähnle	5aef14932a	radeonsi: fix BFE/BFI lowering for GLSL semantics Fixes spec/arb_gpu_shader5/execution/built-in-functions/*-bitfield{Extract,Insert} Cc: 13.0 <mesa-stable@lists.freedesktop.org> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-02 12:30:11 +01:00
Nicolai Hähnle	6526977306	tgsi: align the definition of BFI & [UI]BFE with GLSL As previously written, these opcodes use the SM5 semantics which is incompatible with GLSL when bits == 0, offset == 32. At some point we may want to add BFI_SM5 etc. opcodes, but all users currently either want (and expect!) the GLSL semantics or don't care. Bitfield inserts are generated by the GLSL lower_instructions and lower_packing_builtins passes with constant bits and offset arguments, so any workaround code that drivers may have to emit to follow GLSL semantics should be optimized away easily for those uses. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-11-02 12:30:07 +01:00
Dave Airlie	9f0726f3e5	radv: expose xlib platform extension I missed this when I added the xlib code, this allows dolphin emu to start and crash later. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Cc: "13.0" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-11-02 10:00:38 +10:00
Lionel Landwerlin	a28db12e21	intel: aubinator: print field values if available Turning this : sampler state 0 Sampler Disable: false Texture Border Color Mode: 0 LOD PreClamp Enable: 1 Base Mip Level: 0.000000 Mip Mode Filter: 0 Mag Mode Filter: 1 Min Mode Filter: 1 Texture LOD Bias: foo Anisotropic Algorithm: 0 into this : sampler state 0 Sampler Disable: false Texture Border Color Mode: 0 (DX10/OGL) LOD PreClamp Enable: 1 (OGL) Base Mip Level: 0.000000 Mip Mode Filter: 0 (NONE) Mag Mode Filter: 1 (LINEAR) Min Mode Filter: 1 (LINEAR) Texture LOD Bias: foo Anisotropic Algorithm: 0 (LEGACY) Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Sirisha Gandikota<sirisha.gandikota@intel.com>	2016-11-01 22:37:56 +00:00
Lionel Landwerlin	74c4c84482	intel: aubinator: load fields values from xml data Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Sirisha Gandikota<sirisha.gandikota@intel.com>	2016-11-01 22:37:52 +00:00
Lionel Landwerlin	c8806eeefc	intel: aubinator: print boolean fields to true with colors Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Sirisha Gandikota<sirisha.gandikota@intel.com>	2016-11-01 22:37:22 +00:00
Marek Olšák	d3244c47ce	amd: fix a typo in PIXEL_PIPE_STAT_RESET definition Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-01 22:33:13 +01:00
Marek Olšák	7786f8c635	gallium/radeon: add enum radeon_micro_mode Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-01 22:33:13 +01:00
Marek Olšák	1a4e0162fc	gallium/radeon: make it clear that DRM 2.x.x fast clear constraint is CIK-only Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-01 22:33:13 +01:00
Marek Olšák	e3697b4be6	gallium/radeon: remove r600_surface::level_info Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-01 22:33:13 +01:00
Marek Olšák	bf4d102ea3	gallium/radeon: add radeon_surf::is_linear Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-01 22:33:13 +01:00
Marek Olšák	e9c76eeeaa	gallium/radeon: remove radeon_surf_level::pitch_bytes Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-01 22:33:13 +01:00
Marek Olšák	c66a550385	gallium/radeon: don't call u_format helpers if we have that info already Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-01 22:33:13 +01:00
Marek Olšák	692f2640ab	gallium/radeon: replace radeon_surf_info::dcc_enabled with num_dcc_levels Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-01 22:33:13 +01:00
Marek Olšák	315eb0acb4	radeonsi: add a driver query for counting CP DMA calls CP DMA calls are synchronous with regard to shaders, but can be made asynchronous if needed. Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-01 22:33:13 +01:00
Marek Olšák	d268b7f95e	radeonsi: add a driver query for shader cache hits This is an 8-month old patch. Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-11-01 22:33:13 +01:00
Marek Olšák	6b309f7368	gbm: set up the interop extension for egl/drm breaking libgbm -> libEGL ABI? Acked-by: Alex Deucher <alexander.deucher@amd.com> Reviewed-by: Emil Velikov <emil.velikov@collabora.com>	2016-11-01 22:33:13 +01:00
Samuel Pitoiset	8bfd65395e	nvc0: do not duplicate similar performance metrics Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Pierre Moreau <pierre.morrow@free.fr>	2016-11-01 19:03:26 +01:00
Jason Ekstrand	c41ec1679f	anv/device: Return DEVICE_LOST if execbuf2 fails This makes more sense than OUT_OF_HOST_MEMORY. Technically, you can recover from a failed execbuf2 but the batch you just submitted didn't fully execute so things are in an ill-defined state. The app doesn't want to continue from that point anyway. Signed-off-by: Jason Ekstrand <jason@jlekstrand.net> Cc: "13.0" <mesa-stable@lists.freedesktop.org> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2016-11-01 07:54:52 -07:00
Antia Puentes	61a8a55f55	i965/gen8: Fix vertex attrib upload for dvec3/4 shader inputs The emission of vertex attributes corresponding to dvec3 and dvec4 vertex shader input variables was not correct when the <size> passed to the VertexAttribL* commands was <= 2. This was because we were using the vertex array size when emitting vertices to decide if we uploaded a 64-bit floating point attribute as 1 slot (128-bits) for sizes 1 and 2, or 2 slots (256-bits) for sizes 3 and 4. This caused problems when mapping the input variables to registers because, for deciding which registers contain the values uploaded for a certain variable, we use the size and type given to the variable in the shader, so we will be assigning 256-bits to dvec3/4 variables, even if we only uploaded 128-bits for them, which happened when the vertex array size was <= 2. The patch uses the shader information to only emit as 128-bits those 64-bit floating point variables that were declared as double or dvec2 in the vertex shader. Dvec3 and dvec4 variables will be always uploaded as 256-bits, independently of the <size> given to the VertexAttribL* command. From the ARB_vertex_attrib_64bit specification: "For the 64-bit double precision types listed in Table X.1, no default attribute values are provided if the values of the vertex attribute variable are specified with fewer components than required for the attribute variable. For example, the fourth component of a variable of type dvec4 will be undefined if specified using VertexAttribL3dv or using a vertex array specified with VertexAttribLPointer and a size of three." We are filling these unspecified components with zeros, which coincidentally is also what the GL44-CTS.vertex_attrib_binding.basic-inputL-case1 expects. v2: Do not use bitcount (Kenneth Graunke) Fixes: GL44-CTS.vertex_attrib_binding.basic-inputL-case1 test Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=97287 Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2016-11-01 09:39:09 +01:00
Dave Airlie	f88ea8c72a	radv: drop some unused cmask info members. These were assigned but never used. Inspired by similiar patch in radeonsi. Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-11-01 15:11:35 +10:00
Lionel Landwerlin	1b88760f85	intel: aubinator: fix printing missing gen option Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2016-10-31 22:03:13 +00:00
Lionel Landwerlin	46d67799a6	intel: aubinator: fix assumptions on amount of required data We require 12 bytes of headers but in some cases we just need 4. Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2016-10-31 22:03:09 +00:00
Lionel Landwerlin	6f05b69572	intel: aubinator: don't print out blocks twice Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2016-10-31 22:02:41 +00:00
Nanley Chery	e9a25e0247	i965: Move gen8_disable_stages to brw_upload_initial_gpu_state 3DSTATE_WM_CHROMAKEY isn't programmed anywhere else. 3DSTATE_WM_HZ_OP is programmed, then cleared by blorp during a HZ op, so repeatedly clearing it after every blorp execution is redundant. Signed-off-by: Nanley Chery <nanley.g.chery@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2016-10-31 13:20:05 -07:00
Nanley Chery	477ea60b68	i965: Program 3DSTATE_AA_LINE_PARAMETERS in upload_invariant_state This packet is non-pipelined and doesn't ever change across emissions. Signed-off-by: Nanley Chery <nanley.g.chery@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2016-10-31 13:20:00 -07:00
Leo Liu	06e3cd6a45	st/omx/dec: disable tunnel for size different case When the video coded size is different from frame size, we need the result buffers are same as coded size, which are not size compatible with encode required size, so that simply use no tunnel for this case instead of frame by frame converting. Signed-off-by: Leo Liu <leo.liu@amd.com> Cc: 13.0 <mesa-stable@lists.freedesktop.org>	2016-10-31 11:45:29 -04:00
Leo Liu	d9b2c4048d	st/omx/dec: result buffers size should match codec decoder size Otherwise fails the check of matching between decoder size and buffers size in kernel. Signed-off-by: Leo Liu <leo.liu@amd.com> Cc: 13.0 <mesa-stable@lists.freedesktop.org>	2016-10-31 11:45:14 -04:00
George Kyriazis	55fb874376	swr: [rasterizer] added EventHandlerFile contructor Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2016-10-31 09:06:29 -05:00
George Kyriazis	0a5811b0f3	swr: [rasterizer core] Frontend dependency work Add frontend dependency concept in the DRAW_CONTEXT, which allows serialization of frontend work if necessary. Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2016-10-31 09:06:21 -05:00
George Kyriazis	06f93d0329	swr: [rasterizer core] Refactor/cleanup backends Used for common code reuse and simplification Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2016-10-31 09:06:15 -05:00
George Kyriazis	78a0a09e48	swr: [rasterizer core] Remove deprecated simd intrinsics Used in abandoned all-or-nothing approach to converting to AVX512 Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2016-10-31 09:06:08 -05:00
George Kyriazis	1a3ed86348	swr: [rasterizer archrast] Add thread tags to event files. This allows the post-processor to easily detect the API thread and to process frame information. The frame information is needed to optimized how data is processed from worker threads. Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2016-10-31 09:05:25 -05:00
Marek Olšák	7a2387c3e0	glsl: use a non-malloc'd storage for short ir_variable names Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-10-31 11:53:38 +01:00
Marek Olšák	21e11b5282	glsl: use the linear allocator in opt_constant_propagation Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-10-31 11:53:38 +01:00
Marek Olšák	565b2c4c4b	glsl: use the linear allocator in opt_copy_propagation Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-10-31 11:53:38 +01:00
Marek Olšák	b6f50e4640	glsl: use the linear allocator in opt_copy_propagation_elements Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-10-31 11:53:38 +01:00
Marek Olšák	9c19dedff0	glsl: use the linear allocator in opt_dead_code_local Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-10-31 11:53:38 +01:00
Marek Olšák	23e373eb4f	glsl: use the linear allocator in glsl_symbol_table no ralloc_free occurences Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-10-31 11:53:38 +01:00
Marek Olšák	a4a93103fb	glsl: use the linear allocator for ast_node and derived classes Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-10-31 11:53:38 +01:00
Marek Olšák	2296bb0967	glsl/lexer: use the linear allocator Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-10-31 11:53:38 +01:00
Marek Olšák	47e1758692	glcpp: use the linear allocator for most objects v2: cosmetic changes Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> (v1) Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> (v1)	2016-10-31 11:53:38 +01:00
Marek Olšák	6608dbf540	ralloc: add a linear allocator as a child node of ralloc v2: remove goto, cosmetic changes Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> (v1) Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-10-31 11:53:38 +01:00
Marek Olšák	acc23b04cf	ralloc: remove memset from ralloc_size only do it in rzalloc_size as it was supposed to be Reviewed-by: Edward O'Callaghan <funfunctor@folklore1984.net> Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com>	2016-10-31 11:53:38 +01:00
Marek Olšák	52d2b28f7f	ralloc: use rzalloc where it's necessary No change in behavior. ralloc_size is equivalent to rzalloc_size. That will change though. Calls not switched to rzalloc_size: - ralloc_vasprintf - glsl_type::name allocation (it's filled with snprintf) - C++ classes where valgrind didn't show uninitialized values I switched most of non-glsl stuff to rzalloc without checking whether it's really needed. Reviewed-by: Edward O'Callaghan <funfunctor@folklore1984.net> Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-10-31 11:53:38 +01:00
Marek Olšák	9454f7c0ef	ralloc: add DECLARE_RZALLOC_CXX_OPERATORS Reviewed-by: Edward O'Callaghan <funfunctor@folklore1984.net> Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> Reviewed-by: Kenneth Graunke <kenneth@whiteacpe.org>	2016-10-31 11:53:38 +01:00
Juha-Pekka Heikkila	3bf6c6c3ad	nir: zero allocated memory where needed Signed-off-by: Marek Olšák <marek.olsak@amd.com>	2016-10-31 11:53:38 +01:00
Juha-Pekka Heikkila	4d4335c81a	i965/fs: fill allocated memory with zeros where needed Signed-off-by: Marek Olšák <marek.olsak@amd.com>	2016-10-31 11:53:38 +01:00
Juha-Pekka Heikkila	5fa41520e4	i965/vec4: zero allocated memory where needed Signed-off-by: Juha-Pekka Heikkila <juhapekka.heikkila@gmail.com> Signed-off-by: Marek Olšák <marek.olsak@amd.com>	2016-10-31 11:53:38 +01:00
Tapani Pälli	e40c5dab5e	glsl/glcpp: initialize all fields of glcpp_parser_t on creation this fixes some of the regressions with "ralloc: remove memset from ralloc_size" Signed-off-by: Tapani Pälli <tapani.palli@intel.com> Signed-off-by: Marek Olšák <marek.olsak@amd.com>	2016-10-31 11:53:38 +01:00
Juha-Pekka Heikkila	6770b17b99	glsl: Fix reading of uninitialized memory Switch to use memory allocations which zero memory for places where needed. v2: modify and rebase on top of Marek's series (Tapani) Signed-off-by: Juha-Pekka Heikkila <juhapekka.heikkila@gmail.com> Signed-off-by: Marek Olšák <marek.olsak@amd.com>	2016-10-31 11:53:38 +01:00
Marek Olšák	f67c5a7ccd	glsl: initialize glsl_struct_field properly don't rely on ralloc doing memset Reviewed-by: Edward O'Callaghan <funfunctor@folklore1984.net> Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> Reviewed-by: Kenneth Graunke <kenneth@whiteacpe.org>	2016-10-31 11:53:38 +01:00
Marek Olšák	330482177c	ralloc: don't memset ralloc_header, clear it manually time GALLIUM_NOOP=1 ./run shaders/private/alien_isolation/ >/dev/null Before (2 takes): real 0m8.734s 0m8.773s user 0m34.232s 0m34.348s sys 0m0.084s 0m0.056s After (2 takes): real 0m8.448s 0m8.463s user 0m33.104s 0m33.160s sys 0m0.088s 0m0.076s Average change in "real" time spent: -3.4% calloc should only do 2 things compared to malloc: - check for overflow of "n * size" - call memset I'm not sure if that explains the difference. v2: clear "parent" and "next" in the caller of add_child. Reviewed-by: Edward O'Callaghan <funfunctor@folklore1984.net> (v1) Tested-by: Edmondo Tommasina <edmondo.tommasina@gmail.com> (v1) Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> (v1)	2016-10-31 11:53:38 +01:00
Serge Martin	cb0879985a	clover: Implement clGetExtensionFunctionAddressForPlatform. Add clGetExtensionFunctionAddressForPlatform (CL 1.2). Reviewed-by: Francisco Jerez <currojerez@riseup.net>	2016-10-30 12:53:03 -07:00
Vedran Miletić	2fba72046d	clover: Introduce CLOVER_EXTRA__OPTIONS environment variables The options specified in the CLOVER_EXTRA_BUILD_OPTIONS shell variable are appended to the options specified by the OpenCL program in the clBuildProgram function call, if any. Analogously, the options specified in the CLOVER_EXTRA_COMPILE_OPTIONS and CLOVER_EXTRA_LINK_OPTIONS variables are appended to the options specified in clCompileProgram and clLinkProgram function calls, respectively. v2: rename to CLOVER_EXTRA_COMPILER_OPTIONS * use debug_get_option * append to linker options as well v3: code cleanups v4: separate CLOVER_EXTRA_LINKER_OPTIONS options v5: * fix documentation typo * use CLOVER_EXTRA_COMPILER_OPTIONS in link stage v6: * separate in CLOVER_EXTRA_{BUILD,COMPILE,LINK}_OPTIONS * append options in cl{Build,Compile,Link}Program Signed-off-by: Vedran Miletić <vedran@miletic.net> Reviewed-by[v1]: Edward O'Callaghan <funfunctor@folklore1984.net> v7 [Francisco Jerez]: Slight simplification. Reviewed-by: Francisco Jerez <currojerez@riseup.net>	2016-10-30 12:45:26 -07:00
Vedran Miletić	e3272865c2	clover: Pass unquoted compiler arguments to Clang OpenCL apps can quote arguments they pass to the OpenCL compiler, most commonly include paths containing spaces. If the Clang OpenCL compiler was called via a shell, the shell would split the arguments with respect to to quotes and then remove quotes before passing the arguments to the compiler. Since we call Clang as a library, we have to split the argument with respect to quotes and then remove quotes before passing the arguments. v2: move to tokenize(), remove throwing of CL_INVALID_COMPILER_OPTIONS v3: simplify parsing logic, use more C++11 v4: restore error throwing, clarify a comment Signed-off-by: Vedran Miletić <vedran@miletic.net> Reviewed-by: Francisco Jerez <currojerez@riseup.net>	2016-10-30 12:14:59 -07:00
Jason Ekstrand	2a4a86862c	i965/fs/generator: Don't use the address immediate for MOV_INDIRECT The address immediate field is only 9 bits and, since the value is in bytes, the highest GRF we can point to with it is g15. This makes it pretty close to useless for MOV_INDIRECT. There were already piles of restrictions preventing us from using it prior to Broadwell, so let's get rid of the gen8+ code path entirely. Signed-off-by: Jason Ekstrand <jason@jlekstrand.net> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=97779 Cc: "12.0 13.0" <mesa-stable@lists.freedesktop.org> Reviewed-by: Matt Turner <mattst88@gmail.com>	2016-10-28 17:11:16 -07:00

... 2 3 4 5 6 ...

79281 Commits